Test control evidence samples against the control description
For: Internal control manager or second-line tester running the annual control testing cycle
The pain today
Every key control needs sampled evidence checked: was the approval there, by the right person, before the event, for the right amount. It is repetitive reading of tickets, screenshots and sign-offs, and tired testers pass what should fail.
The ask
“I attached our control descriptions with their test steps and the evidence collected for each sample. For every sample, tell me whether the evidence shows the control operated as described, which test step fails if not, and where evidence is missing or does not match the sample.”
Plain words, as you would say it to a colleague. Edit it to fit your case before you send it.
What you attach or connect
- Control descriptions with test steps and attributes
- Sample lists per control
- Evidence files: tickets, approvals, reports, screenshots
- Delegation of authority and approver lists
The unit of work
One worker task per one evidence sample for one control.
Why a swarm fits
A sample is tested from its own evidence and one control's test steps. Samples are independent, numerous and similar, and each verdict can be checked attribute by attribute.
Not for
Automated controls proven by system configuration, or a framework with a few controls and small samples.
The decision tree
6 typed decisions, each with an action for every answer
At fixed moments in a run, the engine puts one narrow question to a decision model. The decision model never writes text: it answers yes or no with a probability, picks from listed options, or gives a score, about a small slice of the material. The engine then does exactly what this tree says, which is what makes the run auditable. The thresholds are the template's design values, not measured results.
Planner, while planning
Scope checkYes or no, with a probability
Before work starts on a unit
Does the control description name the approver role, the timing and the threshold, so that attributes can be tested?
Sees only: One control description with its test steps
Why: Vague controls are reported as untestable instead of being passed.
- Yes: 0.60 or higherthenAccept
- Unsure: 0.30 up to 0.60thenEscalate to a strong model
- No: below 0.30thenMark unresolved
Scope checkYes or no, with a probability
Before work starts on a unit
Does this evidence file identify the sampled item by its reference, date or amount, rather than a similar transaction?
Sees only: The sample line and the evidence file's identifying fields
Why: Evidence for the wrong item is the commonest silent failure in testing.
- Yes: 0.80 or higherthenAccept
- Unsure: 0.50 up to 0.80thenMark unresolved
- No: below 0.50thenMark unresolved
Before workers, before a task runs
Small worker or strong modelYes or no, with a probability
Before a task runs
Is the evidence a structured ticket or system report, as opposed to a screenshot, a scanned signature or an email chain?
Sees only: The evidence file's type and first page
Why: Messy evidence gets a stronger reader.
- Yes: 0.60 or higherthenAccept
- Unsure: 0.40 up to 0.60thenEscalate to a strong model
- No: below 0.40thenEscalate to a strong model
After workers, the judge checks
Evidence checkYes or no, with a probability
After a worker answers
Does the evidence show approval by a person on the authority list valid that day, dated before the event it approves?
Sees only: The quoted approval, the event date and the authority list extract
Why: A pass must survive re-performance by an auditor.
- Yes: 0.90 or higherthenAccept
- Unsure: 0.50 up to 0.90thenEscalate to a strong model
- No: below 0.50thenReject and retry
Reconciler, while merging
Conflict checkA choice among options
While reconciling
When an attribute fails, what does the evidence show?
Sees only: The failed attribute and the evidence quoted for it
Why: Separates real control failures from evidence-collection gaps.
- Control not performed, or performed latethenAccept
- Evidence missing, control may have runthenMark unresolved
- Wrong evidence attachedthenReject and retry
Accountable person, before anything is settled
Person decidesYes or no, with a probability
Before anything is reported as settled
Would this result be recorded as an exception or deficiency against a named control owner?
Sees only: The test sheet for one sample
Why: The tester concludes on effectiveness and agrees it with the owner.
Accountable: The control tester concludes on each control's effectiveness and agrees deficiencies with the control owner.
- Yes: 0.40 or higherthenAsk a person
- Unsure: 0.15 up to 0.40thenAsk a person
- No: below 0.15thenAccept
The fleet: who does what
Model tiers by role, not brands: you choose the models. Strong reasoning models plan and reconcile, small fast models do the wide work, and the judge is a decision model from a different family, so it does not share the workers' blind spots.
Planner
A strong reasoning model turns each control description into explicit attributes to test and the evidence each one needs.
Decisions here:1. Scope check2. Scope check
Workers
Small fast workers from an open-weight family test one sample, attribute by attribute, quoting the evidence.
Designed for 20 to 800 agents, one worker task per one evidence sample for one control. Each worker receives only its own unit.
Decisions here:3. Small worker or strong model
Judge, from a different model family
A decision model from a different family re-decides every 'pass' on approver, timing and amount against the quoted evidence.
Decisions here:4. Evidence check
Reconciler
A strong reasoning model aggregates exceptions per control and separates evidence gaps from real control failures.
Decisions here:5. Conflict check
Accountable person
The control tester concludes on each control's effectiveness and agrees deficiencies with the control owner.
Decisions here:6. Person decides
Checked before anything is accepted
- Dates are compared in code: approval before event, review within its period
- Approvers are checked against the authority list valid on that date
- Evidence must identify the sampled item, not a similar one
- A pass without quoted evidence for every attribute is refused
What comes back
- Test sheet per sample with attribute results and evidence
- Exceptions per control, with cause
- Samples with missing or mismatched evidence
- Controls whose description is too vague to test
What to measure
- Passes overturned on re-performance by a tester or auditor
- Exceptions found that manual testing missed in the same samples
- Tester hours per sample
Names of measures only. No result is claimed for this template.
Templates open in the workspace chat with the ask filled in. Nothing runs until you send it.
Get early accessSign in to useMore in Compliance and risk
Re-screen counterparties when sanctions lists or ownership change
For: Sanctions or financial crime compliance officer at a trading, shipping or manufacturing group
Screening happens at onboarding and then goes stale.
Map a new regulation to policies and controls, rule by rule
For: Head of compliance or regulatory change manager implementing a new rule
A new regulation arrives with obligations buried in articles, annexes and guidance.
Review vendor assurance packs by security, privacy and resilience
For: Third-party risk manager onboarding or re-assessing critical vendors
Each vendor sends an assurance report, a questionnaire, policies and a penetration test summary.