Support conversation quality review at full coverage

For: Support quality lead responsible for coaching and compliance across the team

Pattern: Specialist panelNeeds scaleDesigned for 40 to 1,200 agents

The pain today

Quality review reads a small random sample. The risky conversations, a wrong answer, a missed identity check, a privacy slip, are mostly in the part nobody reads.

The ask

I attached last month's support conversations, our quality rubric, the help articles and our identity-check and privacy rules. Review every conversation for accuracy, policy, privacy and tone, and show me the ones that need a person, with the message that caused the flag.

Plain words, as you would say it to a colleague. Edit it to fit your case before you send it.

What you attach or connect

  • Conversation export for the period
  • Quality rubric
  • Help articles as the source of correct answers
  • Identity-check and privacy rules

The unit of work

One worker task per one conversation reviewed under one brief.

Why a swarm fits

Each conversation is short and self-contained, and each reviewer holds only its own rulebook. Four narrow briefs across every conversation is wide, shallow work.

Not for

Reviewing one agent's week: a team lead reading those chats gives better coaching.

The decision tree

6 typed decisions, each with an action for every answer

At fixed moments in a run, the engine puts one narrow question to a decision model. The decision model never writes text: it answers yes or no with a probability, picks from listed options, or gives a score, about a small slice of the material. The engine then does exactly what this tree says, which is what makes the run auditable. The thresholds are the template's design values, not measured results.

  1. Before workers, before a task runs

    Small worker or strong modelYes or no, with a probability

    Before a task runs

    Does the conversation involve account access, payment details, an address change or a data request, so that the privacy and identity-check reviewer must read it?

    Sees only: The conversation's topic tags and first messages

    Why: The privacy brief runs only where there is something to check.

    • Yes: 0.50 or higherthenAccept
    • Unsure: 0.20 up to 0.50thenAccept
    • No: below 0.20thenSkip this unit
  2. Small worker or strong modelYes or no, with a probability

    Before a task runs

    Does the agent state a product fact, a limit or a procedure that can be compared with a help article?

    Sees only: The agent's messages in one conversation

    Why: Accuracy review is skipped for conversations that contain nothing to verify.

    • Yes: 0.55 or higherthenAccept
    • Unsure: 0.25 up to 0.55thenAccept
    • No: below 0.25thenSkip this unit
  3. After workers, the judge checks

    Evidence checkYes or no, with a probability

    After a worker answers

    Does the quoted agent message disclose or change account details before the identity-check steps in the quoted rule were completed?

    Sees only: The conversation up to the quoted message and the quoted identity-check rule

    Why: A privacy flag against a named agent must be right before anyone sees it.

    • Yes: 0.90 or higherthenAccept
    • Unsure: 0.50 up to 0.90thenEscalate to a strong model
    • No: below 0.50thenReject and retry
  4. Evidence checkYes or no, with a probability

    After a worker answers

    Does the agent's quoted answer contradict the quoted help article on the same point, rather than add detail the article lacks?

    Sees only: The quoted agent answer and the quoted help article passage

    Why: Separates wrong answers from answers that are merely fuller than the article.

    • Yes: 0.85 or higherthenAccept
    • Unsure: 0.50 up to 0.85thenEscalate to a strong model
    • No: below 0.50thenReject and retry
  5. Run control, between rounds

    Another round?Yes or no, with a probability

    Between rounds

    Did the re-review of unflagged conversations find a violation that the first reviewers missed?

    Sees only: The re-review results for the sampled unflagged conversations

    Why: A miss in the sample widens the re-review; a clean sample ends it.

    • Yes: 0.50 or higherthenContinue
    • Unsure: 0.20 up to 0.50thenContinue
    • No: below 0.20thenStop
  6. Accountable person, before anything is settled

    Person decidesYes or no, with a probability

    Before anything is reported as settled

    Does this flag connect a named employee with a policy or privacy violation?

    Sees only: One verified flag

    Why: Coaching and any consequence are the quality lead's decisions under staff-monitoring rules.

    Accountable: The quality lead decides coaching and any consequence for a named employee, under the employer's staff-monitoring rules.

    • Yes: 0.30 or higherthenAsk a person
    • Unsure: 0.10 up to 0.30thenAsk a person
    • No: below 0.10thenAccept

The fleet: who does what

Model tiers by role, not brands: you choose the models. Strong reasoning models plan and reconcile, small fast models do the wide work, and the judge is a decision model from a different family, so it does not share the workers' blind spots.

  1. Planner

    A strong reasoning model turns rubric and rules into four briefs and routes conversations by channel and topic.

  2. Workers

    Small fast reviewers from mixed open-weight families, one brief each: accuracy, policy, privacy, tone.

    Designed for 40 to 1,200 agents, one worker task per one conversation reviewed under one brief. Each worker receives only its own unit.

    Decisions here:1. Small worker or strong model2. Small worker or strong model

  3. Judge, from a different model family

    A decision model from a different family checks each flag against the quoted message and the quoted rule.

    Decisions here:3. Evidence check4. Evidence check

  4. Reconciler

    A strong reasoning model builds coaching packs and keeps reviewer disagreements on the same conversation visible.

    Decisions here:5. Another round?

  5. Accountable person

    The quality lead decides coaching and any consequence for a named employee, under the employer's staff-monitoring rules.

    Decisions here:6. Person decides

Checked before anything is accepted

  • Each flag quotes the message and the rule it breaks
  • Accuracy flags cite the help article the answer contradicts
  • A sample of unflagged conversations is re-reviewed by another family
  • Tone flags are advisory and never counted as violations

What comes back

  • Conversations needing a person, with the quoted message
  • Coaching packs by theme
  • Help articles that agents contradict repeatedly
  • Reviewer disagreements

What to measure

  • Flags the quality lead upholds
  • Issues found in the unflagged sample
  • Reviewer hours per coaching session
  • Cost per conversation reviewed

Names of measures only. No result is claimed for this template.

Templates open in the workspace chat with the ask filled in. Nothing runs until you send it.

Get early accessSign in to use

A quarter of support tickets clustered into defects

For: Head of support or product operations lead reporting product problems to engineering

Tags are applied in a hurry and mean different things to different agents.

Pattern: Map, verify, reduceNeeds scale5 decisionsDesigned for 50 to 1,000 agents

Help articles re-checked when the product changes

For: Knowledge manager or support enablement lead owning the help centre

Every release quietly breaks a few help articles.

Pattern: WatchtowerNeeds a connector5 decisionsDesigned for 10 to 300 agents

Renewal risk review across the whole customer book

For: Head of customer success or renewals manager preparing the quarter's renewal plan

Health scores are a colour in a dashboard.

Pattern: Map, verify, reduceNeeds a connector5 decisionsDesigned for 30 to 600 agents