Red-team a regulatory impact assessment before it is published

For: Policy adviser, regulatory economist or scrutiny body reviewer

Pattern: Adversarial reviewRuns todayDesigned for 4 to 150 agents

The pain today

An impact assessment makes many factual and causal claims and cites evidence for each. Reviewers rarely have time to open every cited source, so weak claims are found by opponents after publication.

The ask

I attached our draft impact assessment and the studies, statistics and consultation evidence it cites. Split it into its individual claims. For each, show the evidence that supports it, then argue against it using the same sources. Tell me which claims hold, which overstate their source, and which have no support in what I attached.

Plain words, as you would say it to a colleague. Edit it to fit your case before you send it.

What you attach or connect

  • Draft impact assessment as text
  • Cited studies, statistics and reports as text
  • Consultation evidence summary

The unit of work

One worker task per one claim with its cited sources.

Why a swarm fits

A claim is tested against the few sources it cites, so each check is small and independent. An opposing team per claim does what a hostile reader will do after publication.

Not for

Checking the economic model or its calculations, or judging the policy choice. It tests claims against cited text only.

The decision tree

6 typed decisions, each with an action for every answer

At fixed moments in a run, the engine puts one narrow question to a decision model. The decision model never writes text: it answers yes or no with a probability, picks from listed options, or gives a score, about a small slice of the material. The engine then does exactly what this tree says, which is what makes the run auditable. The thresholds are the template's design values, not measured results.

  1. Planner, while planning

    Scope checkYes or no, with a probability

    Before work starts on a unit

    Does this sentence assert a fact, a quantity or a cause and effect, rather than describe the policy or the document's structure?

    Sees only: One sentence of the assessment with its paragraph

    Why: Keeps the claim register to statements that can be tested.

    • Yes: 0.60 or higherthenAccept
    • Unsure: 0.30 up to 0.60thenEscalate to a strong model
    • No: below 0.30thenSkip this unit
  2. Scope checkA choice among options

    Before work starts on a unit

    Is the source cited for this claim among the supplied documents?

    Sees only: The claim's citation and the supplied documents index

    Why: Separates unverifiable from unsupported, which are different problems for the author.

    • SuppliedthenAccept
    • Cited but not suppliedthenMark unresolved
    • No source citedthenMark unresolved
  3. After workers, the judge checks

    Evidence checkYes or no, with a probability

    After a worker answers

    Does the quoted source passage state the claim's figure or finding, without a qualifier the claim leaves out?

    Sees only: One claim and the quoted source passage

    Why: Tests the supporting case the way a hostile reader would after publication.

    • Yes: 0.85 or higherthenAccept
    • Unsure: 0.50 up to 0.85thenEscalate to a strong model
    • No: below 0.50thenReject and retry
  4. Evidence checkA choice among options

    After a worker answers

    Does the source's population, place and period match those the claim applies it to?

    Sees only: One claim and the source's scope statement

    Why: Overstated scope is the most common weakness and is fixable with narrower wording.

    • Same scopethenAccept
    • Source is narrower than the claimthenMark unresolved
    • Source is about a different place or periodthenMark unresolved
    • Scope not stated in the sourcethenEscalate to a strong model
  5. Run control, between rounds

    Another round?Yes or no, with a probability

    Between rounds

    Did the latest challenge round change the status of no claim?

    Sees only: Claim statuses before and after the round

    Why: Ends the red team once further attack changes nothing.

    • Yes: 0.70 or higherthenStop
    • Unsure: 0.40 up to 0.70thenContinue
    • No: below 0.40thenContinue
  6. Accountable person, before anything is settled

    Person decidesYes or no, with a probability

    Before anything is reported as settled

    Is the claim overstated, unsupported or unverifiable, or do several other claims depend on it?

    Sees only: One claim register entry with quotes

    Why: The policy adviser and chief economist own every change and the decision to publish.

    Accountable: The responsible policy adviser and chief economist own every change to the assessment and the decision to publish.

    • Yes: 0.40 or higherthenAsk a person
    • Unsure: 0.15 up to 0.40thenAsk a person
    • No: below 0.15thenAccept

The fleet: who does what

Model tiers by role, not brands: you choose the models. Strong reasoning models plan and reconcile, small fast models do the wide work, and the judge is a decision model from a different family, so it does not share the workers' blind spots.

  1. Planner

    A strong reasoning model splits the assessment into atomic claims and links each to the sources it cites.

    Decisions here:1. Scope check2. Scope check

  2. Workers

    Small fast workers from an open-weight family build the supporting case for one claim with quoted source passages.

    Designed for 4 to 150 agents, one worker task per one claim with its cited sources. Each worker receives only its own unit.

  3. Judge, from a different model family

    Challengers and a decision model from a different family test whether the source says as much as the claim does.

    Decisions here:3. Evidence check4. Evidence check

  4. Reconciler

    A strong reasoning model sorts claims into supported, overstated and unsupported and keeps disputed ones open.

    Decisions here:5. Another round?

  5. Accountable person

    The responsible policy adviser and chief economist own every change to the assessment and the decision to publish.

    Decisions here:6. Person decides

Checked before anything is accepted

  • Each supported claim quotes the source passage it rests on
  • Scope, population and date of the source are compared with those of the claim
  • Claims citing sources that were not supplied are marked unverifiable, not unsupported

What comes back

  • Claim register with status and quotes
  • Overstated claims with the narrower wording the source allows
  • Claims with no cited or supplied support
  • Assumptions that several claims depend on

What to measure

  • Challenges the policy adviser accepts
  • Weak claims later raised by external scrutiny
  • Reviewer time per claim
  • Cost per assessment

Names of measures only. No result is claimed for this template.

Templates open in the workspace chat with the ask filled in. Nothing runs until you send it.

Get early accessSign in to use

Code consultation responses by question and keep minority views

For: Policy analyst or consultation lead at a ministry, regulator or municipality

A public consultation brings thousands of free-text responses, many from campaigns, some from experts with one decisive point.

Pattern: Map, verify, reduceNeeds scale6 decisionsDesigned for 40 to 800 agents

Prepare tender evaluation evidence with two independent readings

For: Procurement officer or evaluation panel chair in a public contracting authority

Evaluators must score long submissions against published criteria and defend every score if challenged.

Pattern: Cross-examinationNeeds live models6 decisionsDesigned for 6 to 300 agents

Trace a changed legal definition through the statute book

For: Legislative drafter, government lawyer or policy adviser preparing an amending bill

Changing one defined term can ripple through acts, regulations and guidance that use or cross-refer to it.

Pattern: Hierarchical decompositionNeeds a connector6 decisionsDesigned for 30 to 600 agents