A quarter of support tickets clustered into defects
For: Head of support or product operations lead reporting product problems to engineering
The pain today
Tags are applied in a hurry and mean different things to different agents. The quarterly top-issues slide is built from tag counts, and engineering does not trust it because nobody can show the tickets.
The ask
“I attached an export of last quarter's tickets with the full conversations. Group them into the actual product problems behind them, not by our tags. For each problem give me the cited tickets, how customers describe it, and which ones look like the same defect worded differently.”
Plain words, as you would say it to a colleague. Edit it to fit your case before you send it.
What you attach or connect
- Ticket export with full conversations
- Product area list and known-issue list
- Release notes for the quarter
- Customer tier per ticket, if available
The unit of work
One worker task per one ticket conversation.
Why a swarm fits
Each ticket is read alone and reduced to a small structured claim: product area, symptom, trigger. Clustering then runs on the claims, so no model ever holds the whole quarter.
Not for
A single angry escalation: a person should read that thread from start to finish.
The decision tree
5 typed decisions, each with an action for every answer
At fixed moments in a run, the engine puts one narrow question to a decision model. The decision model never writes text: it answers yes or no with a probability, picks from listed options, or gives a score, about a small slice of the material. The engine then does exactly what this tree says, which is what makes the run auditable. The thresholds are the template's design values, not measured results.
Planner, while planning
Scope checkA choice among options
Before work starts on a unit
What is this ticket mainly about?
Sees only: The first customer messages of one ticket
Why: Defect clustering should not pay to read how-to questions, but doubtful tickets stay in.
- A product problem or unexpected behaviourthenAccept
- A how-to, billing or account questionthenSkip this unit
- Spam, auto-reply or emptythenSkip this unit
- Cannot tell from the conversationthenAccept
Before workers, before a task runs
Small worker or strong modelYes or no, with a probability
Before a task runs
Is the conversation in one language, with the customer describing the symptom in the first few messages?
Sees only: The ticket's first messages and its length
Why: Long, mixed-language threads are where small workers lose the symptom.
- Yes: 0.60 or higherthenAccept
- Unsure: 0.30 up to 0.60thenEscalate to a strong model
- No: below 0.30thenEscalate to a strong model
After workers, the judge checks
Evidence checkYes or no, with a probability
After a worker answers
Do the customer's quoted words describe the symptom and trigger the worker recorded, without the worker adding a cause the customer never mentioned?
Sees only: The worker's filled form and the quoted customer sentences
Why: Engineering only trusts clusters built from what customers actually said.
- Yes: 0.85 or higherthenAccept
- Unsure: 0.50 up to 0.85thenEscalate to a strong model
- No: below 0.50thenReject and retry
Reconciler, while merging
Conflict checkYes or no, with a probability
While reconciling
Do these two ticket claims describe the same symptom under the same trigger in the same product area, only worded differently?
Sees only: Two filled forms with their quotes
Why: Merging too eagerly invents big defects; doubtful merges stay split.
- Yes: 0.85 or higherthenAccept
- Unsure: 0.50 up to 0.85thenMark unresolved
- No: below 0.50thenContinue
Accountable person, before anything is settled
Person decidesYes or no, with a probability
Before anything is reported as settled
Does this cluster claim a defect that is not on the known-issue list, or do its quotes still contain unmasked personal data?
Sees only: One cluster summary with its quotes and the known-issue list
Why: Product decides what is a real defect, and customer data must not leak into a report.
Accountable: Product and engineering decide what gets fixed and in what order. Ticket text is customer data and needs masking and access limits.
- Yes: 0.30 or higherthenAsk a person
- Unsure: 0.10 up to 0.30thenAsk a person
- No: below 0.10thenAccept
The fleet: who does what
Model tiers by role, not brands: you choose the models. Strong reasoning models plan and reconcile, small fast models do the wide work, and the judge is a decision model from a different family, so it does not share the workers' blind spots.
Planner
A strong reasoning model defines the extraction form and product areas, and revises them after a first sample.
Decisions here:1. Scope check
Workers
Small fast workers from an open-weight family each read one ticket and fill the form with quoted customer words.
Designed for 50 to 1,000 agents, one worker task per one ticket conversation. Each worker receives only its own unit.
Decisions here:2. Small worker or strong model
Judge, from a different model family
A decision model from a different family checks each extracted symptom is supported by the quoted text.
Decisions here:3. Evidence check
Reconciler
A strong reasoning model merges claims into defect clusters and keeps doubtful merges as separate candidates.
Decisions here:4. Conflict check
Accountable person
Product and engineering decide what gets fixed and in what order. Ticket text is customer data and needs masking and access limits.
Decisions here:5. Person decides
Checked before anything is accepted
- Every cluster lists its ticket ids and a quote per ticket
- Tickets that fit no cluster stay in an unassigned pile
- Cluster merges are checked by the judge; doubtful merges stay split
- Personal data in quotes is masked before anything is reported
What comes back
- Defect clusters with cited tickets and customer wording
- Links to known issues and releases
- New problems not on the known-issue list
- Unassigned tickets
- How the tags compare with the clusters
What to measure
- Clusters engineering accepts as real defects
- Tickets wrongly placed, from a sampled check
- Analyst hours to the quarterly report
- Cost per ticket read
Names of measures only. No result is claimed for this template.
Templates open in the workspace chat with the ask filled in. Nothing runs until you send it.
Get early accessSign in to useMore in Customer support and success
Help articles re-checked when the product changes
For: Knowledge manager or support enablement lead owning the help centre
Every release quietly breaks a few help articles.
Renewal risk review across the whole customer book
For: Head of customer success or renewals manager preparing the quarter's renewal plan
Health scores are a colour in a dashboard.
Reply drafts narrowed against policy and help articles
For: Support operations lead writing macros for the most common contact reasons
Macros are written once by whoever had time.