Dual-screen abstracts for a systematic review

For: Evidence synthesis lead, health economist or medical affairs researcher running a systematic review

Pattern: Cross-examinationRuns todayDesigned for 3 to 200 agents

The pain today

Review methods call for two independent screeners per abstract. Finding two people for thousands of abstracts is the slow step, and reasons for exclusion are recorded thinly.

The ask

I attached our review protocol with the inclusion and exclusion criteria and the abstracts from our searches as text. Screen each abstract twice, independently, against the criteria. Tell me include, exclude or unclear, quote the sentence that decided it, and list every abstract where the two screens disagree.

Plain words, as you would say it to a colleague. Edit it to fit your case before you send it.

What you attach or connect

  • Review protocol with eligibility criteria
  • Titles and abstracts exported as text

The unit of work

One worker task per one abstract.

Why a swarm fits

An abstract is judged alone against a short list of criteria. The method itself demands two independent readers, which maps directly to two model families that do not share errors.

Not for

Full-text appraisal, risk-of-bias judgement or data extraction from figures. A review with a few dozen abstracts does not need it.

The decision tree

5 typed decisions, each with an action for every answer

At fixed moments in a run, the engine puts one narrow question to a decision model. The decision model never writes text: it answers yes or no with a probability, picks from listed options, or gives a score, about a small slice of the material. The engine then does exactly what this tree says, which is what makes the run auditable. The thresholds are the template's design values, not measured results.

  1. Planner, while planning

    Scope checkYes or no, with a probability

    Before work starts on a unit

    Does this record contain an abstract with enough text to describe a study, rather than a title only, an erratum or a conference listing?

    Sees only: One exported record

    Why: Marks thin records as unclear early instead of letting a screener guess.

    • Yes: 0.60 or higherthenAccept
    • Unsure: 0.30 up to 0.60thenEscalate to a strong model
    • No: below 0.30thenSkip this unit
  2. After workers, the judge checks

    Evidence checkA choice among options

    After a worker answers

    Does the quoted sentence state the study design, population or intervention that the named exclusion criterion rules out?

    Sees only: One quoted abstract sentence and one exclusion criterion

    Why: Keeps exclusions tied to words in the abstract, so no study is dropped on an assumption.

    • States an excluded featurethenAccept
    • States an eligible featurethenReject and retry
    • Does not address this criterionthenMark unresolved
  3. Reconciler, while merging

    Conflict checkA choice among options

    While reconciling

    Did the two independent screens reach the same decision on this abstract?

    Sees only: Both screening decisions with their reasons and quotes

    Why: Never auto-resolves a split toward exclusion; splits go to the human adjudicators.

    • Both includethenAccept
    • Both exclude, same criterionthenAccept
    • Both exclude, different criteriathenEscalate to a strong model
    • One includes, one excludesthenMark unresolved
  4. Run control, between rounds

    Another round?Yes or no, with a probability

    Between rounds

    On the latest calibration batch, do the screens' decisions match the human reviewers' decisions on every included study?

    Sees only: The calibration batch decisions side by side

    Why: Stops the wide pass if the criteria as written are losing eligible studies.

    • Yes: 0.70 or higherthenStop
    • Unsure: 0.40 up to 0.70thenContinue
    • No: below 0.40thenContinue
  5. Accountable person, before anything is settled

    Person decidesYes or no, with a probability

    Before anything is reported as settled

    Is this abstract marked unclear, split between the screens, or excluded on a criterion the protocol lists as needing judgement?

    Sees only: One screening record with both decisions

    Why: Hands every doubtful abstract to the review lead, who owns the included set.

    Accountable: The review lead adjudicates every disagreement and owns the final included set and the methods statement.

    • Yes: 0.40 or higherthenAsk a person
    • Unsure: 0.15 up to 0.40thenAsk a person
    • No: below 0.15thenAccept

The fleet: who does what

Model tiers by role, not brands: you choose the models. Strong reasoning models plan and reconcile, small fast models do the wide work, and the judge is a decision model from a different family, so it does not share the workers' blind spots.

  1. Planner

    A strong reasoning model turns the protocol into explicit screening questions and an order for applying them.

    Decisions here:1. Scope check

  2. Workers

    Two sets of small workers from different families each screen the same abstract without seeing the other's call.

    Designed for 3 to 200 agents, one worker task per one abstract. Each worker receives only its own unit.

  3. Judge, from a different model family

    A decision model from a third family checks that the quoted sentence supports the stated exclusion reason.

    Decisions here:2. Evidence check

  4. Reconciler

    A strong reasoning model tallies decisions and hands every disagreement and every unclear call to the human reviewers.

    Decisions here:3. Conflict check4. Another round?

  5. Accountable person

    The review lead adjudicates every disagreement and owns the final included set and the methods statement.

    Decisions here:5. Person decides

Checked before anything is accepted

  • Every exclusion names a criterion and quotes the abstract
  • Disagreements between the two screens are never auto-resolved to exclude
  • Abstracts with too little information are marked unclear

What comes back

  • Screening table with both decisions, reasons and quotes
  • Disagreement list for human adjudication
  • Exclusion reasons tallied for the flow diagram

What to measure

  • Agreement with human screeners on a calibration sample
  • Included studies the swarm would have excluded
  • Share of abstracts sent to humans
  • Cost per abstract

Names of measures only. No result is claimed for this template.

Templates open in the workspace chat with the ask filled in. Nothing runs until you send it.

Get early accessSign in to use

Grade protocol deviations across trial sites, with evidence

For: Clinical operations lead or trial quality manager at a sponsor or contract research organisation

Deviation logs and monitoring reports pile up across sites.

Pattern: Map, verify, reduceNeeds scale6 decisionsDesigned for 40 to 400 agents

Check adverse-event narratives against their case data

For: Pharmacovigilance lead or safety writer preparing case narratives for a study report or aggregate report

Hundreds of case narratives are written from line listings.

Pattern: Cross-examinationNeeds scale5 decisionsDesigned for 30 to 300 agents