Draft and narrow assessment questions per learning objective
For: Instructional designer or assessment lead building a question bank for a course
The pain today
Writing good questions is slow, so banks are thin and reused. Weak items are ambiguous, test recall of a phrase, or cannot be answered from the course material at all.
The ask
“I attached the course modules and the list of learning objectives. For each objective, draft several candidate questions with answers, then narrow them to the ones that are answerable from the module text, unambiguous and test the objective. Show the source passage for each.”
Plain words, as you would say it to a colleague. Edit it to fit your case before you send it.
What you attach or connect
- Course module texts
- Learning objectives
- Item writing guideline
- Existing question bank, to avoid duplicates
The unit of work
One worker task per one learning objective and its module text.
Why a swarm fits
Candidates are cheap to draft from one objective and a few pages. Judging is separate work: independent judges answer each item blind from the text, which exposes ambiguous or unanswerable ones.
Not for
High-stakes exams without psychometric review. Item difficulty and bias need piloting with real learners.
The decision tree
6 typed decisions, each with an action for every answer
At fixed moments in a run, the engine puts one narrow question to a decision model. The decision model never writes text: it answers yes or no with a probability, picks from listed options, or gives a score, about a small slice of the material. The engine then does exactly what this tree says, which is what makes the run auditable. The thresholds are the template's design values, not measured results.
Planner, while planning
Scope checkYes or no, with a probability
Before work starts on a unit
Do the paired module passages contain enough content to write a question that tests this learning objective?
Sees only: One learning objective and the module passages paired with it
Why: An objective the material never covers is a finding for the course team, not a drafting task.
- Yes: 0.60 or higherthenAccept
- Unsure: 0.30 up to 0.60thenAccept
- No: below 0.30thenMark unresolved
After workers, the judge checks
Evidence checkYes or no, with a probability
After a worker answers
Does the source passage make the keyed answer correct and each of the other options incorrect?
Sees only: One candidate item with its key and options, and its source passage
Why: An item with two defensible answers is unfair to learners and is the most common drafting fault.
- Yes: 0.85 or higherthenAccept
- Unsure: 0.50 up to 0.85thenEscalate to a strong model
- No: below 0.50thenReject and retry
Evidence checkA choice among options
After a worker answers
Which option does the module passage support when the item is read without its answer key?
Sees only: One candidate item without its key, and the module passage
Why: Answering blind exposes ambiguity that a judge shown the key would overlook.
- The keyed option onlythenAccept
- A different optionthenReject and retry
- More than one optionthenReject and retry
- None from this passagethenSkip this unit
Before workers, before a task runs
Small worker or strong modelA score
Before a task runs
How well does answering this item require the skill the learning objective names, rather than recall of a phrase from the text?
Sees only: One learning objective and one candidate item that passed the blind check
Why: Narrows each round to items that test the objective, and sends borderline ones to a stronger judge.
- High: 0.70 or higherthenAccept
- Middle: 0.40 up to 0.70thenEscalate to a strong model
- Low: below 0.40thenSkip this unit
Run control, between rounds
Another round?Yes or no, with a probability
Between rounds
Does this objective still have fewer surviving items than the item writing guideline asks for?
Sees only: One objective's surviving item count and the guideline's stated requirement
Why: Drafts more only where the bank is still thin.
- Yes: 0.60 or higherthenContinue
- Unsure: 0.30 up to 0.60thenStop
- No: below 0.30thenStop
Accountable person, before anything is settled
Person decidesYes or no, with a probability
Before anything is reported as settled
Is this surviving item intended for use with learners in an assessment that counts?
Sees only: One surviving item with its key, passage and round history
Why: Every item is approved by the assessment lead, who owns fairness and accessibility review.
Accountable: The assessment lead approves every item before use and owns the fairness and accessibility review.
- Yes: 0.10 or higherthenAsk a person
- Unsure: 0.02 up to 0.10thenAsk a person
- No: below 0.02thenAsk a person
The fleet: who does what
Model tiers by role, not brands: you choose the models. Strong reasoning models plan and reconcile, small fast models do the wide work, and the judge is a decision model from a different family, so it does not share the workers' blind spots.
Planner
A strong reasoning model pairs each objective with its module passages and sets the rounds and criteria.
Decisions here:1. Scope check
Workers
Small fast workers from an open-weight family draft candidate items with answer keys and source passages.
Designed for 20 to 500 agents, one worker task per one learning objective and its module text. Each worker receives only its own unit.
Decisions here:4. Small worker or strong model
Judge, from a different model family
Decision models from a different family answer each item blind and score it: answerable, unambiguous, on objective.
Decisions here:2. Evidence check3. Evidence check
Reconciler
A strong reasoning model keeps the surviving items per objective and logs why the others were dropped.
Decisions here:5. Another round?
Accountable person
The assessment lead approves every item before use and owns the fairness and accessibility review.
Decisions here:6. Person decides
Checked before anything is accepted
- An item survives only if blind judges reach the keyed answer from the module text
- Each item links to the passage that makes its answer correct
- Items that can be answered without reading the module are dropped
- Near-duplicates of existing bank items are removed
What comes back
- Surviving items per objective, with keys and source passages
- Dropped items and the reason, per round
- Objectives for which no good item survived
- Passages in the module that the judges found ambiguous
What to measure
- Share of surviving items the assessment lead approves
- Items withdrawn after learner use
- Designer hours per approved item
- Cost per approved item
Names of measures only. No result is claimed for this template.
Templates open in the workspace chat with the ask filled in. Nothing runs until you send it.
Get early accessSign in to useMore in Education and publishing
Pre-review a manuscript for method, statistics and ethics
For: Journal managing editor or handling editor triaging submissions before peer review
Reviewers are scarce.
Check that every citation supports its sentence
For: Production editor or research integrity officer at a publisher or university press
A monograph or report carries hundreds of citations.
Map course modules to accreditation learning outcomes
For: Programme director or quality assurance lead preparing an accreditation review
Accreditors ask where each required outcome is taught and assessed.