Red-team a regulatory impact assessment before it is published
For: Policy adviser, regulatory economist or scrutiny body reviewer
The pain today
An impact assessment makes many factual and causal claims and cites evidence for each. Reviewers rarely have time to open every cited source, so weak claims are found by opponents after publication.
The ask
“I attached our draft impact assessment and the studies, statistics and consultation evidence it cites. Split it into its individual claims. For each, show the evidence that supports it, then argue against it using the same sources. Tell me which claims hold, which overstate their source, and which have no support in what I attached.”
Plain words, as you would say it to a colleague. Edit it to fit your case before you send it.
What you attach or connect
- Draft impact assessment as text
- Cited studies, statistics and reports as text
- Consultation evidence summary
The unit of work
One worker task per one claim with its cited sources.
Why a swarm fits
A claim is tested against the few sources it cites, so each check is small and independent. An opposing team per claim does what a hostile reader will do after publication.
Not for
Checking the economic model or its calculations, or judging the policy choice. It tests claims against cited text only.
The decision tree
6 typed decisions, each with an action for every answer
At fixed moments in a run, the engine puts one narrow question to a decision model. The decision model never writes text: it answers yes or no with a probability, picks from listed options, or gives a score, about a small slice of the material. The engine then does exactly what this tree says, which is what makes the run auditable. The thresholds are the template's design values, not measured results.
Planner, while planning
Scope checkYes or no, with a probability
Before work starts on a unit
Does this sentence assert a fact, a quantity or a cause and effect, rather than describe the policy or the document's structure?
Sees only: One sentence of the assessment with its paragraph
Why: Keeps the claim register to statements that can be tested.
- Yes: 0.60 or higherthenAccept
- Unsure: 0.30 up to 0.60thenEscalate to a strong model
- No: below 0.30thenSkip this unit
Scope checkA choice among options
Before work starts on a unit
Is the source cited for this claim among the supplied documents?
Sees only: The claim's citation and the supplied documents index
Why: Separates unverifiable from unsupported, which are different problems for the author.
- SuppliedthenAccept
- Cited but not suppliedthenMark unresolved
- No source citedthenMark unresolved
After workers, the judge checks
Evidence checkYes or no, with a probability
After a worker answers
Does the quoted source passage state the claim's figure or finding, without a qualifier the claim leaves out?
Sees only: One claim and the quoted source passage
Why: Tests the supporting case the way a hostile reader would after publication.
- Yes: 0.85 or higherthenAccept
- Unsure: 0.50 up to 0.85thenEscalate to a strong model
- No: below 0.50thenReject and retry
Evidence checkA choice among options
After a worker answers
Does the source's population, place and period match those the claim applies it to?
Sees only: One claim and the source's scope statement
Why: Overstated scope is the most common weakness and is fixable with narrower wording.
- Same scopethenAccept
- Source is narrower than the claimthenMark unresolved
- Source is about a different place or periodthenMark unresolved
- Scope not stated in the sourcethenEscalate to a strong model
Run control, between rounds
Another round?Yes or no, with a probability
Between rounds
Did the latest challenge round change the status of no claim?
Sees only: Claim statuses before and after the round
Why: Ends the red team once further attack changes nothing.
- Yes: 0.70 or higherthenStop
- Unsure: 0.40 up to 0.70thenContinue
- No: below 0.40thenContinue
Accountable person, before anything is settled
Person decidesYes or no, with a probability
Before anything is reported as settled
Is the claim overstated, unsupported or unverifiable, or do several other claims depend on it?
Sees only: One claim register entry with quotes
Why: The policy adviser and chief economist own every change and the decision to publish.
Accountable: The responsible policy adviser and chief economist own every change to the assessment and the decision to publish.
- Yes: 0.40 or higherthenAsk a person
- Unsure: 0.15 up to 0.40thenAsk a person
- No: below 0.15thenAccept
The fleet: who does what
Model tiers by role, not brands: you choose the models. Strong reasoning models plan and reconcile, small fast models do the wide work, and the judge is a decision model from a different family, so it does not share the workers' blind spots.
Planner
A strong reasoning model splits the assessment into atomic claims and links each to the sources it cites.
Decisions here:1. Scope check2. Scope check
Workers
Small fast workers from an open-weight family build the supporting case for one claim with quoted source passages.
Designed for 4 to 150 agents, one worker task per one claim with its cited sources. Each worker receives only its own unit.
Judge, from a different model family
Challengers and a decision model from a different family test whether the source says as much as the claim does.
Decisions here:3. Evidence check4. Evidence check
Reconciler
A strong reasoning model sorts claims into supported, overstated and unsupported and keeps disputed ones open.
Decisions here:5. Another round?
Accountable person
The responsible policy adviser and chief economist own every change to the assessment and the decision to publish.
Decisions here:6. Person decides
Checked before anything is accepted
- Each supported claim quotes the source passage it rests on
- Scope, population and date of the source are compared with those of the claim
- Claims citing sources that were not supplied are marked unverifiable, not unsupported
What comes back
- Claim register with status and quotes
- Overstated claims with the narrower wording the source allows
- Claims with no cited or supplied support
- Assumptions that several claims depend on
What to measure
- Challenges the policy adviser accepts
- Weak claims later raised by external scrutiny
- Reviewer time per claim
- Cost per assessment
Names of measures only. No result is claimed for this template.
Templates open in the workspace chat with the ask filled in. Nothing runs until you send it.
Get early accessSign in to useMore in Public sector and policy
Code consultation responses by question and keep minority views
For: Policy analyst or consultation lead at a ministry, regulator or municipality
A public consultation brings thousands of free-text responses, many from campaigns, some from experts with one decisive point.
Prepare tender evaluation evidence with two independent readings
For: Procurement officer or evaluation panel chair in a public contracting authority
Evaluators must score long submissions against published criteria and defend every score if challenged.
Trace a changed legal definition through the statute book
For: Legislative drafter, government lawyer or policy adviser preparing an amending bill
Changing one defined term can ripple through acts, regulations and guidance that use or cross-refer to it.