Dual-screen abstracts for a systematic review
For: Evidence synthesis lead, health economist or medical affairs researcher running a systematic review
The pain today
Review methods call for two independent screeners per abstract. Finding two people for thousands of abstracts is the slow step, and reasons for exclusion are recorded thinly.
The ask
“I attached our review protocol with the inclusion and exclusion criteria and the abstracts from our searches as text. Screen each abstract twice, independently, against the criteria. Tell me include, exclude or unclear, quote the sentence that decided it, and list every abstract where the two screens disagree.”
Plain words, as you would say it to a colleague. Edit it to fit your case before you send it.
What you attach or connect
- Review protocol with eligibility criteria
- Titles and abstracts exported as text
The unit of work
One worker task per one abstract.
Why a swarm fits
An abstract is judged alone against a short list of criteria. The method itself demands two independent readers, which maps directly to two model families that do not share errors.
Not for
Full-text appraisal, risk-of-bias judgement or data extraction from figures. A review with a few dozen abstracts does not need it.
The decision tree
5 typed decisions, each with an action for every answer
At fixed moments in a run, the engine puts one narrow question to a decision model. The decision model never writes text: it answers yes or no with a probability, picks from listed options, or gives a score, about a small slice of the material. The engine then does exactly what this tree says, which is what makes the run auditable. The thresholds are the template's design values, not measured results.
Planner, while planning
Scope checkYes or no, with a probability
Before work starts on a unit
Does this record contain an abstract with enough text to describe a study, rather than a title only, an erratum or a conference listing?
Sees only: One exported record
Why: Marks thin records as unclear early instead of letting a screener guess.
- Yes: 0.60 or higherthenAccept
- Unsure: 0.30 up to 0.60thenEscalate to a strong model
- No: below 0.30thenSkip this unit
After workers, the judge checks
Evidence checkA choice among options
After a worker answers
Does the quoted sentence state the study design, population or intervention that the named exclusion criterion rules out?
Sees only: One quoted abstract sentence and one exclusion criterion
Why: Keeps exclusions tied to words in the abstract, so no study is dropped on an assumption.
- States an excluded featurethenAccept
- States an eligible featurethenReject and retry
- Does not address this criterionthenMark unresolved
Reconciler, while merging
Conflict checkA choice among options
While reconciling
Did the two independent screens reach the same decision on this abstract?
Sees only: Both screening decisions with their reasons and quotes
Why: Never auto-resolves a split toward exclusion; splits go to the human adjudicators.
- Both includethenAccept
- Both exclude, same criterionthenAccept
- Both exclude, different criteriathenEscalate to a strong model
- One includes, one excludesthenMark unresolved
Run control, between rounds
Another round?Yes or no, with a probability
Between rounds
On the latest calibration batch, do the screens' decisions match the human reviewers' decisions on every included study?
Sees only: The calibration batch decisions side by side
Why: Stops the wide pass if the criteria as written are losing eligible studies.
- Yes: 0.70 or higherthenStop
- Unsure: 0.40 up to 0.70thenContinue
- No: below 0.40thenContinue
Accountable person, before anything is settled
Person decidesYes or no, with a probability
Before anything is reported as settled
Is this abstract marked unclear, split between the screens, or excluded on a criterion the protocol lists as needing judgement?
Sees only: One screening record with both decisions
Why: Hands every doubtful abstract to the review lead, who owns the included set.
Accountable: The review lead adjudicates every disagreement and owns the final included set and the methods statement.
- Yes: 0.40 or higherthenAsk a person
- Unsure: 0.15 up to 0.40thenAsk a person
- No: below 0.15thenAccept
The fleet: who does what
Model tiers by role, not brands: you choose the models. Strong reasoning models plan and reconcile, small fast models do the wide work, and the judge is a decision model from a different family, so it does not share the workers' blind spots.
Planner
A strong reasoning model turns the protocol into explicit screening questions and an order for applying them.
Decisions here:1. Scope check
Workers
Two sets of small workers from different families each screen the same abstract without seeing the other's call.
Designed for 3 to 200 agents, one worker task per one abstract. Each worker receives only its own unit.
Judge, from a different model family
A decision model from a third family checks that the quoted sentence supports the stated exclusion reason.
Decisions here:2. Evidence check
Reconciler
A strong reasoning model tallies decisions and hands every disagreement and every unclear call to the human reviewers.
Decisions here:3. Conflict check4. Another round?
Accountable person
The review lead adjudicates every disagreement and owns the final included set and the methods statement.
Decisions here:5. Person decides
Checked before anything is accepted
- Every exclusion names a criterion and quotes the abstract
- Disagreements between the two screens are never auto-resolved to exclude
- Abstracts with too little information are marked unclear
What comes back
- Screening table with both decisions, reasons and quotes
- Disagreement list for human adjudication
- Exclusion reasons tallied for the flow diagram
What to measure
- Agreement with human screeners on a calibration sample
- Included studies the swarm would have excluded
- Share of abstracts sent to humans
- Cost per abstract
Names of measures only. No result is claimed for this template.
Templates open in the workspace chat with the ask filled in. Nothing runs until you send it.
Get early accessSign in to useMore in Healthcare and life sciences
Grade protocol deviations across trial sites, with evidence
For: Clinical operations lead or trial quality manager at a sponsor or contract research organisation
Deviation logs and monitoring reports pile up across sites.
Check adverse-event narratives against their case data
For: Pharmacovigilance lead or safety writer preparing case narratives for a study report or aggregate report
Hundreds of case narratives are written from line listings.
Trace every summary claim in a dossier back to its study report
For: Regulatory affairs lead or medical writer assembling a marketing application
Summary documents restate results from many study reports.