Compare root causes across outage and fault reports

For: Reliability engineer or head of operations at a network operator or plant fleet owner

Pattern: Adversarial reviewRuns todayDesigned for 4 to 120 agents

The pain today

Each outage gets its own investigation and its own root cause label. Nobody tests whether the stated cause holds against the report's own timeline, or whether one weakness keeps returning under different names.

The ask

I attached our outage investigation reports. Group them by real underlying cause, not by the label each report used. Then try to break each grouping using the reports' own timelines and protection sequences. Tell me which common causes survive and which reports state a cause their own evidence does not support.

Plain words, as you would say it to a colleague. Edit it to fit your case before you send it.

What you attach or connect

  • Outage and fault investigation reports as text
  • Event timelines and protection operation sequences as text
  • Corrective action lists per event

The unit of work

One worker task per one outage report.

Why a swarm fits

Reports are read independently, then a second team attacks each proposed common cause with the same reports. A single reader tends to accept the label on the cover.

Not for

A single event investigation. It works from written reports and does not read fault recorder traces, relay event files or control system logs.

The decision tree

6 typed decisions, each with an action for every answer

At fixed moments in a run, the engine puts one narrow question to a decision model. The decision model never writes text: it answers yes or no with a probability, picks from listed options, or gives a score, about a small slice of the material. The engine then does exactly what this tree says, which is what makes the run auditable. The thresholds are the template's design values, not measured results.

  1. Planner, while planning

    Scope checkYes or no, with a probability

    Before work starts on a unit

    Does this report contain a dated event sequence and a stated root cause?

    Sees only: The report's contents page and findings section

    Why: Reports without a sequence cannot be tested and are listed rather than grouped.

    • Yes: 0.60 or higherthenAccept
    • Unsure: 0.30 up to 0.60thenEscalate to a strong model
    • No: below 0.30thenSkip this unit
  2. After workers, the judge checks

    Evidence checkYes or no, with a probability

    After a worker answers

    Does the quoted passage from the report's findings state the underlying cause the worker proposed, in the report's own words?

    Sees only: The proposed cause and one quoted findings passage

    Why: Keeps the building team from inventing causes the investigators never wrote.

    • Yes: 0.85 or higherthenAccept
    • Unsure: 0.50 up to 0.85thenEscalate to a strong model
    • No: below 0.50thenReject and retry
  3. Evidence checkA choice among options

    After a worker answers

    Does the quoted event sequence show the stated cause occurring before the first protection operation?

    Sees only: The stated cause and the quoted event sequence

    Why: Lets the challengers test each cause against the report's own timeline.

    • Sequence places it beforethenAccept
    • Sequence places it after or not at allthenMark unresolved
    • Timestamps missingthenMark unresolved
  4. Reconciler, while merging

    Conflict checkA choice among options

    While reconciling

    Do these two reports describe the same failed component type and the same failure mechanism, whatever label each used?

    Sees only: The cause passages of two reports

    Why: Builds groupings from what reports describe, not from their cover labels.

    • Same component and mechanismthenAccept
    • Same component, different mechanismthenMark unresolved
    • Different componentthenContinue
  5. Run control, between rounds

    Another round?Yes or no, with a probability

    Between rounds

    Did the latest challenge round leave every grouping's membership unchanged?

    Sees only: Grouping membership before and after the round

    Why: Stops paying for challenge rounds once the groupings have settled.

    • Yes: 0.70 or higherthenStop
    • Unsure: 0.40 up to 0.70thenContinue
    • No: below 0.40thenContinue
  6. Accountable person, before anything is settled

    Person decidesYes or no, with a probability

    Before anything is reported as settled

    Does the finding dispute a report's stated root cause, or involve a protection operation or a repeated corrective action?

    Sees only: One grouping or contested report with quotes

    Why: The reliability engineer owns every root-cause conclusion; the swarm only shows where the paper disagrees.

    Accountable: The reliability engineer owns the root-cause finding; the operations head owns any change to settings, maintenance policy or procedure.

    • Yes: 0.40 or higherthenAsk a person
    • Unsure: 0.15 up to 0.40thenAsk a person
    • No: below 0.15thenAccept

The fleet: who does what

Model tiers by role, not brands: you choose the models. Strong reasoning models plan and reconcile, small fast models do the wide work, and the judge is a decision model from a different family, so it does not share the workers' blind spots.

  1. Planner

    A strong reasoning model sets the cause taxonomy and assigns reports to a building team and a challenging team.

    Decisions here:1. Scope check

  2. Workers

    Small fast workers from an open-weight family summarise one report's sequence and propose its underlying cause with quotes.

    Designed for 4 to 120 agents, one worker task per one outage report. Each worker receives only its own unit.

  3. Judge, from a different model family

    Challengers and a decision model from a different family test each grouping against timelines and reject unsupported links.

    Decisions here:2. Evidence check3. Evidence check

  4. Reconciler

    A strong reasoning model keeps the groupings that survive and lists the contested ones with both arguments.

    Decisions here:4. Conflict check5. Another round?

  5. Accountable person

    The reliability engineer owns the root-cause finding; the operations head owns any change to settings, maintenance policy or procedure.

    Decisions here:6. Person decides

Checked before anything is accepted

  • Each cause is tied to quoted passages from the report's timeline or findings
  • A grouping is kept only if the challenge fails on every member report
  • Reports whose stated cause conflicts with their own sequence are listed separately

What comes back

  • Surviving common causes with member reports and quotes
  • Contested groupings with both arguments
  • Reports whose stated root cause lacks support
  • Corrective actions repeated across events

What to measure

  • Groupings the reliability engineer accepts
  • Recurring causes not previously tracked
  • Engineer hours per report reviewed
  • Cost per report

Names of measures only. No result is claimed for this template.

Templates open in the workspace chat with the ask filled in. Nothing runs until you send it.

Get early accessSign in to use

Reconcile study assumptions across an interconnection queue

For: Interconnection manager at a developer, or a transmission planner reviewing clustered studies

Each project's study rests on a base case, on earlier-queued projects assumed in service, on short-circuit levels and on shared network upgrades.

Pattern: Map, verify, reduceNeeds scale6 decisionsDesigned for 30 to 400 agents

Check a flexibility contract portfolio against dispatch obligations

For: Portfolio manager at a flexibility aggregator, or a utility demand-response programme lead

Every site contract has its own notice time, duration limit, activation cap, availability window, baseline method and penalty.

Pattern: Specialist panelNeeds live models6 decisionsDesigned for 8 to 300 agents

Review maintenance and test records across a substation fleet

For: Asset manager or maintenance engineer at a transmission or distribution network operator

Oil analysis reports, breaker timing tests, relay test sheets and battery discharge tests sit inside work orders.

Pattern: Hierarchical decompositionNeeds a connector6 decisionsDesigned for 40 to 600 agents