Reply drafts narrowed against policy and help articles
For: Support operations lead writing macros for the most common contact reasons
The pain today
Macros are written once by whoever had time. Some promise things the policy does not allow, some are cold, and testing alternatives by hand never reaches the top of the list.
The ask
“I attached our refund and warranty policy, the related help articles, our tone guide and example tickets for our most common contact reasons. Write many reply drafts for each reason, remove any that promise something the policy does not say, and give me the best few with the policy line they rest on.”
Plain words, as you would say it to a colleague. Edit it to fit your case before you send it.
What you attach or connect
- Refund, warranty and service policy
- Related help articles
- Tone of voice guide
- Example tickets per contact reason
The unit of work
One worker task per one candidate reply for one contact reason.
Why a swarm fits
Drafts are cheap and independent. The costly part is judging them, and a judge needs only the draft, the policy passage and the tone guide.
Not for
A reply to one customer: an agent with the policy open writes it faster.
The decision tree
5 typed decisions, each with an action for every answer
At fixed moments in a run, the engine puts one narrow question to a decision model. The decision model never writes text: it answers yes or no with a probability, picks from listed options, or gives a score, about a small slice of the material. The engine then does exactly what this tree says, which is what makes the run auditable. The thresholds are the template's design values, not measured results.
After workers, the judge checks
Evidence checkYes or no, with a probability
After a worker answers
Does every promise in this draft, such as a refund window or a free replacement, appear in the quoted policy line with the same conditions?
Sees only: One draft reply and the policy passage for its contact reason
Why: A macro is sent many times; one promise the policy does not make is repeated every time.
- Yes: 0.90 or higherthenAccept
- Unsure: 0.50 up to 0.90thenEscalate to a strong model
- No: below 0.50thenSkip this unit
Evidence checkA score
After a worker answers
How clearly would a customer with this contact reason know what happens next, and what they must do, after reading this draft?
Sees only: One draft reply, an example ticket for the reason and the tone guide
Why: Narrows on usefulness to the customer instead of on polish.
- High: 0.70 or higherthenContinue
- Middle: 0.40 up to 0.70thenContinue
- Low: below 0.40thenSkip this unit
Reconciler, while merging
Conflict checkYes or no, with a probability
While reconciling
Do these two judges' verdicts on the same draft differ on whether it is accurate to the policy?
Sees only: Two judges' score sheets for one draft
Why: Disagreement about accuracy is never averaged; a stronger judge settles it or the owner sees it.
- Yes: 0.60 or higherthenEscalate to a strong model
- Unsure: 0.30 up to 0.60thenEscalate to a strong model
- No: below 0.30thenAccept
Run control, between rounds
Another round?Yes or no, with a probability
Between rounds
Did the last round change which drafts lead for any contact reason?
Sees only: The ranking per contact reason before and after the round
Why: Ends the tournament once further rounds only reshuffle the same winners.
- Yes: 0.60 or higherthenContinue
- Unsure: 0.30 up to 0.60thenStop
- No: below 0.30thenStop
Accountable person, before anything is settled
Person decidesYes or no, with a probability
Before anything is reported as settled
Is there no policy line covering what customers with this contact reason ask for, so that any reply would in effect set policy?
Sees only: The contact reason, its example tickets and the policy index
Why: Policy gaps belong to the policy owner, not to a model or a macro.
Accountable: The support lead approves macros before use; policy gaps go to whoever owns the policy, not to a model.
- Yes: 0.35 or higherthenAsk a person
- Unsure: 0.10 up to 0.35thenAsk a person
- No: below 0.10thenAccept
The fleet: who does what
Model tiers by role, not brands: you choose the models. Strong reasoning models plan and reconcile, small fast models do the wide work, and the judge is a decision model from a different family, so it does not share the workers' blind spots.
Planner
A strong reasoning model lists contact reasons, the policy passages governing each, and the scoring rubric.
Workers
Small fast workers from mixed open-weight families each write drafts for one contact reason in varied styles.
Designed for 8 to 120 agents, one worker task per one candidate reply for one contact reason. Each worker receives only its own unit.
Judge, from a different model family
Judges from a different family than the authors score policy accuracy first, then clarity and tone, in rounds.
Decisions here:1. Evidence check2. Evidence check
Reconciler
A strong reasoning model removes near-duplicates and presents survivors with scores and open judge disagreements.
Decisions here:3. Conflict check4. Another round?
Accountable person
The support lead approves macros before use; policy gaps go to whoever owns the policy, not to a model.
Decisions here:5. Person decides
Checked before anything is accepted
- Any promise in a draft must match a quoted policy line or the draft is removed
- Judges never score drafts from their own model family
- Judge disagreements in the final round are shown to the owner
What comes back
- Short list per contact reason with the policy citation
- Removed drafts with the policy reason
- Contact reasons where the policy is silent
What to measure
- Drafts adopted as macros
- Policy errors later found in adopted macros
- Customer satisfaction on tickets using the new macros
- Cost per adopted macro
Names of measures only. No result is claimed for this template.
Templates open in the workspace chat with the ask filled in. Nothing runs until you send it.
Get early accessSign in to useMore in Customer support and success
A quarter of support tickets clustered into defects
For: Head of support or product operations lead reporting product problems to engineering
Tags are applied in a hurry and mean different things to different agents.
Help articles re-checked when the product changes
For: Knowledge manager or support enablement lead owning the help centre
Every release quietly breaks a few help articles.
Renewal risk review across the whole customer book
For: Head of customer success or renewals manager preparing the quarter's renewal plan
Health scores are a colour in a dashboard.