Test control evidence samples against the control description

For: Internal control manager or second-line tester running the annual control testing cycle

Pattern: Map, verify, reduceNeeds scaleDesigned for 20 to 800 agents

The pain today

Every key control needs sampled evidence checked: was the approval there, by the right person, before the event, for the right amount. It is repetitive reading of tickets, screenshots and sign-offs, and tired testers pass what should fail.

The ask

I attached our control descriptions with their test steps and the evidence collected for each sample. For every sample, tell me whether the evidence shows the control operated as described, which test step fails if not, and where evidence is missing or does not match the sample.

Plain words, as you would say it to a colleague. Edit it to fit your case before you send it.

What you attach or connect

  • Control descriptions with test steps and attributes
  • Sample lists per control
  • Evidence files: tickets, approvals, reports, screenshots
  • Delegation of authority and approver lists

The unit of work

One worker task per one evidence sample for one control.

Why a swarm fits

A sample is tested from its own evidence and one control's test steps. Samples are independent, numerous and similar, and each verdict can be checked attribute by attribute.

Not for

Automated controls proven by system configuration, or a framework with a few controls and small samples.

The decision tree

6 typed decisions, each with an action for every answer

At fixed moments in a run, the engine puts one narrow question to a decision model. The decision model never writes text: it answers yes or no with a probability, picks from listed options, or gives a score, about a small slice of the material. The engine then does exactly what this tree says, which is what makes the run auditable. The thresholds are the template's design values, not measured results.

  1. Planner, while planning

    Scope checkYes or no, with a probability

    Before work starts on a unit

    Does the control description name the approver role, the timing and the threshold, so that attributes can be tested?

    Sees only: One control description with its test steps

    Why: Vague controls are reported as untestable instead of being passed.

    • Yes: 0.60 or higherthenAccept
    • Unsure: 0.30 up to 0.60thenEscalate to a strong model
    • No: below 0.30thenMark unresolved
  2. Scope checkYes or no, with a probability

    Before work starts on a unit

    Does this evidence file identify the sampled item by its reference, date or amount, rather than a similar transaction?

    Sees only: The sample line and the evidence file's identifying fields

    Why: Evidence for the wrong item is the commonest silent failure in testing.

    • Yes: 0.80 or higherthenAccept
    • Unsure: 0.50 up to 0.80thenMark unresolved
    • No: below 0.50thenMark unresolved
  3. Before workers, before a task runs

    Small worker or strong modelYes or no, with a probability

    Before a task runs

    Is the evidence a structured ticket or system report, as opposed to a screenshot, a scanned signature or an email chain?

    Sees only: The evidence file's type and first page

    Why: Messy evidence gets a stronger reader.

    • Yes: 0.60 or higherthenAccept
    • Unsure: 0.40 up to 0.60thenEscalate to a strong model
    • No: below 0.40thenEscalate to a strong model
  4. After workers, the judge checks

    Evidence checkYes or no, with a probability

    After a worker answers

    Does the evidence show approval by a person on the authority list valid that day, dated before the event it approves?

    Sees only: The quoted approval, the event date and the authority list extract

    Why: A pass must survive re-performance by an auditor.

    • Yes: 0.90 or higherthenAccept
    • Unsure: 0.50 up to 0.90thenEscalate to a strong model
    • No: below 0.50thenReject and retry
  5. Reconciler, while merging

    Conflict checkA choice among options

    While reconciling

    When an attribute fails, what does the evidence show?

    Sees only: The failed attribute and the evidence quoted for it

    Why: Separates real control failures from evidence-collection gaps.

    • Control not performed, or performed latethenAccept
    • Evidence missing, control may have runthenMark unresolved
    • Wrong evidence attachedthenReject and retry
  6. Accountable person, before anything is settled

    Person decidesYes or no, with a probability

    Before anything is reported as settled

    Would this result be recorded as an exception or deficiency against a named control owner?

    Sees only: The test sheet for one sample

    Why: The tester concludes on effectiveness and agrees it with the owner.

    Accountable: The control tester concludes on each control's effectiveness and agrees deficiencies with the control owner.

    • Yes: 0.40 or higherthenAsk a person
    • Unsure: 0.15 up to 0.40thenAsk a person
    • No: below 0.15thenAccept

The fleet: who does what

Model tiers by role, not brands: you choose the models. Strong reasoning models plan and reconcile, small fast models do the wide work, and the judge is a decision model from a different family, so it does not share the workers' blind spots.

  1. Planner

    A strong reasoning model turns each control description into explicit attributes to test and the evidence each one needs.

    Decisions here:1. Scope check2. Scope check

  2. Workers

    Small fast workers from an open-weight family test one sample, attribute by attribute, quoting the evidence.

    Designed for 20 to 800 agents, one worker task per one evidence sample for one control. Each worker receives only its own unit.

    Decisions here:3. Small worker or strong model

  3. Judge, from a different model family

    A decision model from a different family re-decides every 'pass' on approver, timing and amount against the quoted evidence.

    Decisions here:4. Evidence check

  4. Reconciler

    A strong reasoning model aggregates exceptions per control and separates evidence gaps from real control failures.

    Decisions here:5. Conflict check

  5. Accountable person

    The control tester concludes on each control's effectiveness and agrees deficiencies with the control owner.

    Decisions here:6. Person decides

Checked before anything is accepted

  • Dates are compared in code: approval before event, review within its period
  • Approvers are checked against the authority list valid on that date
  • Evidence must identify the sampled item, not a similar one
  • A pass without quoted evidence for every attribute is refused

What comes back

  • Test sheet per sample with attribute results and evidence
  • Exceptions per control, with cause
  • Samples with missing or mismatched evidence
  • Controls whose description is too vague to test

What to measure

  • Passes overturned on re-performance by a tester or auditor
  • Exceptions found that manual testing missed in the same samples
  • Tester hours per sample

Names of measures only. No result is claimed for this template.

Templates open in the workspace chat with the ask filled in. Nothing runs until you send it.

Get early accessSign in to use

Map a new regulation to policies and controls, rule by rule

For: Head of compliance or regulatory change manager implementing a new rule

A new regulation arrives with obligations buried in articles, annexes and guidance.

Pattern: Hierarchical decompositionNeeds live models6 decisionsDesigned for 8 to 400 agents