Draft and narrow assessment questions per learning objective

For: Instructional designer or assessment lead building a question bank for a course

Pattern: TournamentNeeds live modelsDesigned for 20 to 500 agents

The pain today

Writing good questions is slow, so banks are thin and reused. Weak items are ambiguous, test recall of a phrase, or cannot be answered from the course material at all.

The ask

I attached the course modules and the list of learning objectives. For each objective, draft several candidate questions with answers, then narrow them to the ones that are answerable from the module text, unambiguous and test the objective. Show the source passage for each.

Plain words, as you would say it to a colleague. Edit it to fit your case before you send it.

What you attach or connect

  • Course module texts
  • Learning objectives
  • Item writing guideline
  • Existing question bank, to avoid duplicates

The unit of work

One worker task per one learning objective and its module text.

Why a swarm fits

Candidates are cheap to draft from one objective and a few pages. Judging is separate work: independent judges answer each item blind from the text, which exposes ambiguous or unanswerable ones.

Not for

High-stakes exams without psychometric review. Item difficulty and bias need piloting with real learners.

The decision tree

6 typed decisions, each with an action for every answer

At fixed moments in a run, the engine puts one narrow question to a decision model. The decision model never writes text: it answers yes or no with a probability, picks from listed options, or gives a score, about a small slice of the material. The engine then does exactly what this tree says, which is what makes the run auditable. The thresholds are the template's design values, not measured results.

  1. Planner, while planning

    Scope checkYes or no, with a probability

    Before work starts on a unit

    Do the paired module passages contain enough content to write a question that tests this learning objective?

    Sees only: One learning objective and the module passages paired with it

    Why: An objective the material never covers is a finding for the course team, not a drafting task.

    • Yes: 0.60 or higherthenAccept
    • Unsure: 0.30 up to 0.60thenAccept
    • No: below 0.30thenMark unresolved
  2. After workers, the judge checks

    Evidence checkYes or no, with a probability

    After a worker answers

    Does the source passage make the keyed answer correct and each of the other options incorrect?

    Sees only: One candidate item with its key and options, and its source passage

    Why: An item with two defensible answers is unfair to learners and is the most common drafting fault.

    • Yes: 0.85 or higherthenAccept
    • Unsure: 0.50 up to 0.85thenEscalate to a strong model
    • No: below 0.50thenReject and retry
  3. Evidence checkA choice among options

    After a worker answers

    Which option does the module passage support when the item is read without its answer key?

    Sees only: One candidate item without its key, and the module passage

    Why: Answering blind exposes ambiguity that a judge shown the key would overlook.

    • The keyed option onlythenAccept
    • A different optionthenReject and retry
    • More than one optionthenReject and retry
    • None from this passagethenSkip this unit
  4. Before workers, before a task runs

    Small worker or strong modelA score

    Before a task runs

    How well does answering this item require the skill the learning objective names, rather than recall of a phrase from the text?

    Sees only: One learning objective and one candidate item that passed the blind check

    Why: Narrows each round to items that test the objective, and sends borderline ones to a stronger judge.

    • High: 0.70 or higherthenAccept
    • Middle: 0.40 up to 0.70thenEscalate to a strong model
    • Low: below 0.40thenSkip this unit
  5. Run control, between rounds

    Another round?Yes or no, with a probability

    Between rounds

    Does this objective still have fewer surviving items than the item writing guideline asks for?

    Sees only: One objective's surviving item count and the guideline's stated requirement

    Why: Drafts more only where the bank is still thin.

    • Yes: 0.60 or higherthenContinue
    • Unsure: 0.30 up to 0.60thenStop
    • No: below 0.30thenStop
  6. Accountable person, before anything is settled

    Person decidesYes or no, with a probability

    Before anything is reported as settled

    Is this surviving item intended for use with learners in an assessment that counts?

    Sees only: One surviving item with its key, passage and round history

    Why: Every item is approved by the assessment lead, who owns fairness and accessibility review.

    Accountable: The assessment lead approves every item before use and owns the fairness and accessibility review.

    • Yes: 0.10 or higherthenAsk a person
    • Unsure: 0.02 up to 0.10thenAsk a person
    • No: below 0.02thenAsk a person

The fleet: who does what

Model tiers by role, not brands: you choose the models. Strong reasoning models plan and reconcile, small fast models do the wide work, and the judge is a decision model from a different family, so it does not share the workers' blind spots.

  1. Planner

    A strong reasoning model pairs each objective with its module passages and sets the rounds and criteria.

    Decisions here:1. Scope check

  2. Workers

    Small fast workers from an open-weight family draft candidate items with answer keys and source passages.

    Designed for 20 to 500 agents, one worker task per one learning objective and its module text. Each worker receives only its own unit.

    Decisions here:4. Small worker or strong model

  3. Judge, from a different model family

    Decision models from a different family answer each item blind and score it: answerable, unambiguous, on objective.

    Decisions here:2. Evidence check3. Evidence check

  4. Reconciler

    A strong reasoning model keeps the surviving items per objective and logs why the others were dropped.

    Decisions here:5. Another round?

  5. Accountable person

    The assessment lead approves every item before use and owns the fairness and accessibility review.

    Decisions here:6. Person decides

Checked before anything is accepted

  • An item survives only if blind judges reach the keyed answer from the module text
  • Each item links to the passage that makes its answer correct
  • Items that can be answered without reading the module are dropped
  • Near-duplicates of existing bank items are removed

What comes back

  • Surviving items per objective, with keys and source passages
  • Dropped items and the reason, per round
  • Objectives for which no good item survived
  • Passages in the module that the judges found ambiguous

What to measure

  • Share of surviving items the assessment lead approves
  • Items withdrawn after learner use
  • Designer hours per approved item
  • Cost per approved item

Names of measures only. No result is claimed for this template.

Templates open in the workspace chat with the ask filled in. Nothing runs until you send it.

Get early accessSign in to use

Check that every citation supports its sentence

For: Production editor or research integrity officer at a publisher or university press

A monograph or report carries hundreds of citations.

Pattern: Map, verify, reduceNeeds scale6 decisionsDesigned for 40 to 1,200 agents

Map course modules to accreditation learning outcomes

For: Programme director or quality assurance lead preparing an accreditation review

Accreditors ask where each required outcome is taught and assessed.

Pattern: Hierarchical decompositionNeeds scale6 decisionsDesigned for 30 to 700 agents