A quarter of support tickets clustered into defects

For: Head of support or product operations lead reporting product problems to engineering

Pattern: Map, verify, reduceNeeds scaleDesigned for 50 to 1,000 agents

The pain today

Tags are applied in a hurry and mean different things to different agents. The quarterly top-issues slide is built from tag counts, and engineering does not trust it because nobody can show the tickets.

The ask

I attached an export of last quarter's tickets with the full conversations. Group them into the actual product problems behind them, not by our tags. For each problem give me the cited tickets, how customers describe it, and which ones look like the same defect worded differently.

Plain words, as you would say it to a colleague. Edit it to fit your case before you send it.

What you attach or connect

  • Ticket export with full conversations
  • Product area list and known-issue list
  • Release notes for the quarter
  • Customer tier per ticket, if available

The unit of work

One worker task per one ticket conversation.

Why a swarm fits

Each ticket is read alone and reduced to a small structured claim: product area, symptom, trigger. Clustering then runs on the claims, so no model ever holds the whole quarter.

Not for

A single angry escalation: a person should read that thread from start to finish.

The decision tree

5 typed decisions, each with an action for every answer

At fixed moments in a run, the engine puts one narrow question to a decision model. The decision model never writes text: it answers yes or no with a probability, picks from listed options, or gives a score, about a small slice of the material. The engine then does exactly what this tree says, which is what makes the run auditable. The thresholds are the template's design values, not measured results.

  1. Planner, while planning

    Scope checkA choice among options

    Before work starts on a unit

    What is this ticket mainly about?

    Sees only: The first customer messages of one ticket

    Why: Defect clustering should not pay to read how-to questions, but doubtful tickets stay in.

    • A product problem or unexpected behaviourthenAccept
    • A how-to, billing or account questionthenSkip this unit
    • Spam, auto-reply or emptythenSkip this unit
    • Cannot tell from the conversationthenAccept
  2. Before workers, before a task runs

    Small worker or strong modelYes or no, with a probability

    Before a task runs

    Is the conversation in one language, with the customer describing the symptom in the first few messages?

    Sees only: The ticket's first messages and its length

    Why: Long, mixed-language threads are where small workers lose the symptom.

    • Yes: 0.60 or higherthenAccept
    • Unsure: 0.30 up to 0.60thenEscalate to a strong model
    • No: below 0.30thenEscalate to a strong model
  3. After workers, the judge checks

    Evidence checkYes or no, with a probability

    After a worker answers

    Do the customer's quoted words describe the symptom and trigger the worker recorded, without the worker adding a cause the customer never mentioned?

    Sees only: The worker's filled form and the quoted customer sentences

    Why: Engineering only trusts clusters built from what customers actually said.

    • Yes: 0.85 or higherthenAccept
    • Unsure: 0.50 up to 0.85thenEscalate to a strong model
    • No: below 0.50thenReject and retry
  4. Reconciler, while merging

    Conflict checkYes or no, with a probability

    While reconciling

    Do these two ticket claims describe the same symptom under the same trigger in the same product area, only worded differently?

    Sees only: Two filled forms with their quotes

    Why: Merging too eagerly invents big defects; doubtful merges stay split.

    • Yes: 0.85 or higherthenAccept
    • Unsure: 0.50 up to 0.85thenMark unresolved
    • No: below 0.50thenContinue
  5. Accountable person, before anything is settled

    Person decidesYes or no, with a probability

    Before anything is reported as settled

    Does this cluster claim a defect that is not on the known-issue list, or do its quotes still contain unmasked personal data?

    Sees only: One cluster summary with its quotes and the known-issue list

    Why: Product decides what is a real defect, and customer data must not leak into a report.

    Accountable: Product and engineering decide what gets fixed and in what order. Ticket text is customer data and needs masking and access limits.

    • Yes: 0.30 or higherthenAsk a person
    • Unsure: 0.10 up to 0.30thenAsk a person
    • No: below 0.10thenAccept

The fleet: who does what

Model tiers by role, not brands: you choose the models. Strong reasoning models plan and reconcile, small fast models do the wide work, and the judge is a decision model from a different family, so it does not share the workers' blind spots.

  1. Planner

    A strong reasoning model defines the extraction form and product areas, and revises them after a first sample.

    Decisions here:1. Scope check

  2. Workers

    Small fast workers from an open-weight family each read one ticket and fill the form with quoted customer words.

    Designed for 50 to 1,000 agents, one worker task per one ticket conversation. Each worker receives only its own unit.

    Decisions here:2. Small worker or strong model

  3. Judge, from a different model family

    A decision model from a different family checks each extracted symptom is supported by the quoted text.

    Decisions here:3. Evidence check

  4. Reconciler

    A strong reasoning model merges claims into defect clusters and keeps doubtful merges as separate candidates.

    Decisions here:4. Conflict check

  5. Accountable person

    Product and engineering decide what gets fixed and in what order. Ticket text is customer data and needs masking and access limits.

    Decisions here:5. Person decides

Checked before anything is accepted

  • Every cluster lists its ticket ids and a quote per ticket
  • Tickets that fit no cluster stay in an unassigned pile
  • Cluster merges are checked by the judge; doubtful merges stay split
  • Personal data in quotes is masked before anything is reported

What comes back

  • Defect clusters with cited tickets and customer wording
  • Links to known issues and releases
  • New problems not on the known-issue list
  • Unassigned tickets
  • How the tags compare with the clusters

What to measure

  • Clusters engineering accepts as real defects
  • Tickets wrongly placed, from a sampled check
  • Analyst hours to the quarterly report
  • Cost per ticket read

Names of measures only. No result is claimed for this template.

Templates open in the workspace chat with the ask filled in. Nothing runs until you send it.

Get early accessSign in to use

Help articles re-checked when the product changes

For: Knowledge manager or support enablement lead owning the help centre

Every release quietly breaks a few help articles.

Pattern: WatchtowerNeeds a connector5 decisionsDesigned for 10 to 300 agents

Renewal risk review across the whole customer book

For: Head of customer success or renewals manager preparing the quarter's renewal plan

Health scores are a colour in a dashboard.

Pattern: Map, verify, reduceNeeds a connector5 decisionsDesigned for 30 to 600 agents