Reply drafts narrowed against policy and help articles

For: Support operations lead writing macros for the most common contact reasons

Pattern: TournamentNeeds live modelsDesigned for 8 to 120 agents

The pain today

Macros are written once by whoever had time. Some promise things the policy does not allow, some are cold, and testing alternatives by hand never reaches the top of the list.

The ask

I attached our refund and warranty policy, the related help articles, our tone guide and example tickets for our most common contact reasons. Write many reply drafts for each reason, remove any that promise something the policy does not say, and give me the best few with the policy line they rest on.

Plain words, as you would say it to a colleague. Edit it to fit your case before you send it.

What you attach or connect

  • Refund, warranty and service policy
  • Related help articles
  • Tone of voice guide
  • Example tickets per contact reason

The unit of work

One worker task per one candidate reply for one contact reason.

Why a swarm fits

Drafts are cheap and independent. The costly part is judging them, and a judge needs only the draft, the policy passage and the tone guide.

Not for

A reply to one customer: an agent with the policy open writes it faster.

The decision tree

5 typed decisions, each with an action for every answer

At fixed moments in a run, the engine puts one narrow question to a decision model. The decision model never writes text: it answers yes or no with a probability, picks from listed options, or gives a score, about a small slice of the material. The engine then does exactly what this tree says, which is what makes the run auditable. The thresholds are the template's design values, not measured results.

  1. After workers, the judge checks

    Evidence checkYes or no, with a probability

    After a worker answers

    Does every promise in this draft, such as a refund window or a free replacement, appear in the quoted policy line with the same conditions?

    Sees only: One draft reply and the policy passage for its contact reason

    Why: A macro is sent many times; one promise the policy does not make is repeated every time.

    • Yes: 0.90 or higherthenAccept
    • Unsure: 0.50 up to 0.90thenEscalate to a strong model
    • No: below 0.50thenSkip this unit
  2. Evidence checkA score

    After a worker answers

    How clearly would a customer with this contact reason know what happens next, and what they must do, after reading this draft?

    Sees only: One draft reply, an example ticket for the reason and the tone guide

    Why: Narrows on usefulness to the customer instead of on polish.

    • High: 0.70 or higherthenContinue
    • Middle: 0.40 up to 0.70thenContinue
    • Low: below 0.40thenSkip this unit
  3. Reconciler, while merging

    Conflict checkYes or no, with a probability

    While reconciling

    Do these two judges' verdicts on the same draft differ on whether it is accurate to the policy?

    Sees only: Two judges' score sheets for one draft

    Why: Disagreement about accuracy is never averaged; a stronger judge settles it or the owner sees it.

    • Yes: 0.60 or higherthenEscalate to a strong model
    • Unsure: 0.30 up to 0.60thenEscalate to a strong model
    • No: below 0.30thenAccept
  4. Run control, between rounds

    Another round?Yes or no, with a probability

    Between rounds

    Did the last round change which drafts lead for any contact reason?

    Sees only: The ranking per contact reason before and after the round

    Why: Ends the tournament once further rounds only reshuffle the same winners.

    • Yes: 0.60 or higherthenContinue
    • Unsure: 0.30 up to 0.60thenStop
    • No: below 0.30thenStop
  5. Accountable person, before anything is settled

    Person decidesYes or no, with a probability

    Before anything is reported as settled

    Is there no policy line covering what customers with this contact reason ask for, so that any reply would in effect set policy?

    Sees only: The contact reason, its example tickets and the policy index

    Why: Policy gaps belong to the policy owner, not to a model or a macro.

    Accountable: The support lead approves macros before use; policy gaps go to whoever owns the policy, not to a model.

    • Yes: 0.35 or higherthenAsk a person
    • Unsure: 0.10 up to 0.35thenAsk a person
    • No: below 0.10thenAccept

The fleet: who does what

Model tiers by role, not brands: you choose the models. Strong reasoning models plan and reconcile, small fast models do the wide work, and the judge is a decision model from a different family, so it does not share the workers' blind spots.

  1. Planner

    A strong reasoning model lists contact reasons, the policy passages governing each, and the scoring rubric.

  2. Workers

    Small fast workers from mixed open-weight families each write drafts for one contact reason in varied styles.

    Designed for 8 to 120 agents, one worker task per one candidate reply for one contact reason. Each worker receives only its own unit.

  3. Judge, from a different model family

    Judges from a different family than the authors score policy accuracy first, then clarity and tone, in rounds.

    Decisions here:1. Evidence check2. Evidence check

  4. Reconciler

    A strong reasoning model removes near-duplicates and presents survivors with scores and open judge disagreements.

    Decisions here:3. Conflict check4. Another round?

  5. Accountable person

    The support lead approves macros before use; policy gaps go to whoever owns the policy, not to a model.

    Decisions here:5. Person decides

Checked before anything is accepted

  • Any promise in a draft must match a quoted policy line or the draft is removed
  • Judges never score drafts from their own model family
  • Judge disagreements in the final round are shown to the owner

What comes back

  • Short list per contact reason with the policy citation
  • Removed drafts with the policy reason
  • Contact reasons where the policy is silent

What to measure

  • Drafts adopted as macros
  • Policy errors later found in adopted macros
  • Customer satisfaction on tickets using the new macros
  • Cost per adopted macro

Names of measures only. No result is claimed for this template.

Templates open in the workspace chat with the ask filled in. Nothing runs until you send it.

Get early accessSign in to use

A quarter of support tickets clustered into defects

For: Head of support or product operations lead reporting product problems to engineering

Tags are applied in a hurry and mean different things to different agents.

Pattern: Map, verify, reduceNeeds scale5 decisionsDesigned for 50 to 1,000 agents

Help articles re-checked when the product changes

For: Knowledge manager or support enablement lead owning the help centre

Every release quietly breaks a few help articles.

Pattern: WatchtowerNeeds a connector5 decisionsDesigned for 10 to 300 agents

Renewal risk review across the whole customer book

For: Head of customer success or renewals manager preparing the quarter's renewal plan

Health scores are a colour in a dashboard.

Pattern: Map, verify, reduceNeeds a connector5 decisionsDesigned for 30 to 600 agents