Compare root causes across outage and fault reports
For: Reliability engineer or head of operations at a network operator or plant fleet owner
The pain today
Each outage gets its own investigation and its own root cause label. Nobody tests whether the stated cause holds against the report's own timeline, or whether one weakness keeps returning under different names.
The ask
“I attached our outage investigation reports. Group them by real underlying cause, not by the label each report used. Then try to break each grouping using the reports' own timelines and protection sequences. Tell me which common causes survive and which reports state a cause their own evidence does not support.”
Plain words, as you would say it to a colleague. Edit it to fit your case before you send it.
What you attach or connect
- Outage and fault investigation reports as text
- Event timelines and protection operation sequences as text
- Corrective action lists per event
The unit of work
One worker task per one outage report.
Why a swarm fits
Reports are read independently, then a second team attacks each proposed common cause with the same reports. A single reader tends to accept the label on the cover.
Not for
A single event investigation. It works from written reports and does not read fault recorder traces, relay event files or control system logs.
The decision tree
6 typed decisions, each with an action for every answer
At fixed moments in a run, the engine puts one narrow question to a decision model. The decision model never writes text: it answers yes or no with a probability, picks from listed options, or gives a score, about a small slice of the material. The engine then does exactly what this tree says, which is what makes the run auditable. The thresholds are the template's design values, not measured results.
Planner, while planning
Scope checkYes or no, with a probability
Before work starts on a unit
Does this report contain a dated event sequence and a stated root cause?
Sees only: The report's contents page and findings section
Why: Reports without a sequence cannot be tested and are listed rather than grouped.
- Yes: 0.60 or higherthenAccept
- Unsure: 0.30 up to 0.60thenEscalate to a strong model
- No: below 0.30thenSkip this unit
After workers, the judge checks
Evidence checkYes or no, with a probability
After a worker answers
Does the quoted passage from the report's findings state the underlying cause the worker proposed, in the report's own words?
Sees only: The proposed cause and one quoted findings passage
Why: Keeps the building team from inventing causes the investigators never wrote.
- Yes: 0.85 or higherthenAccept
- Unsure: 0.50 up to 0.85thenEscalate to a strong model
- No: below 0.50thenReject and retry
Evidence checkA choice among options
After a worker answers
Does the quoted event sequence show the stated cause occurring before the first protection operation?
Sees only: The stated cause and the quoted event sequence
Why: Lets the challengers test each cause against the report's own timeline.
- Sequence places it beforethenAccept
- Sequence places it after or not at allthenMark unresolved
- Timestamps missingthenMark unresolved
Reconciler, while merging
Conflict checkA choice among options
While reconciling
Do these two reports describe the same failed component type and the same failure mechanism, whatever label each used?
Sees only: The cause passages of two reports
Why: Builds groupings from what reports describe, not from their cover labels.
- Same component and mechanismthenAccept
- Same component, different mechanismthenMark unresolved
- Different componentthenContinue
Run control, between rounds
Another round?Yes or no, with a probability
Between rounds
Did the latest challenge round leave every grouping's membership unchanged?
Sees only: Grouping membership before and after the round
Why: Stops paying for challenge rounds once the groupings have settled.
- Yes: 0.70 or higherthenStop
- Unsure: 0.40 up to 0.70thenContinue
- No: below 0.40thenContinue
Accountable person, before anything is settled
Person decidesYes or no, with a probability
Before anything is reported as settled
Does the finding dispute a report's stated root cause, or involve a protection operation or a repeated corrective action?
Sees only: One grouping or contested report with quotes
Why: The reliability engineer owns every root-cause conclusion; the swarm only shows where the paper disagrees.
Accountable: The reliability engineer owns the root-cause finding; the operations head owns any change to settings, maintenance policy or procedure.
- Yes: 0.40 or higherthenAsk a person
- Unsure: 0.15 up to 0.40thenAsk a person
- No: below 0.15thenAccept
The fleet: who does what
Model tiers by role, not brands: you choose the models. Strong reasoning models plan and reconcile, small fast models do the wide work, and the judge is a decision model from a different family, so it does not share the workers' blind spots.
Planner
A strong reasoning model sets the cause taxonomy and assigns reports to a building team and a challenging team.
Decisions here:1. Scope check
Workers
Small fast workers from an open-weight family summarise one report's sequence and propose its underlying cause with quotes.
Designed for 4 to 120 agents, one worker task per one outage report. Each worker receives only its own unit.
Judge, from a different model family
Challengers and a decision model from a different family test each grouping against timelines and reject unsupported links.
Decisions here:2. Evidence check3. Evidence check
Reconciler
A strong reasoning model keeps the groupings that survive and lists the contested ones with both arguments.
Decisions here:4. Conflict check5. Another round?
Accountable person
The reliability engineer owns the root-cause finding; the operations head owns any change to settings, maintenance policy or procedure.
Decisions here:6. Person decides
Checked before anything is accepted
- Each cause is tied to quoted passages from the report's timeline or findings
- A grouping is kept only if the challenge fails on every member report
- Reports whose stated cause conflicts with their own sequence are listed separately
What comes back
- Surviving common causes with member reports and quotes
- Contested groupings with both arguments
- Reports whose stated root cause lacks support
- Corrective actions repeated across events
What to measure
- Groupings the reliability engineer accepts
- Recurring causes not previously tracked
- Engineer hours per report reviewed
- Cost per report
Names of measures only. No result is claimed for this template.
Templates open in the workspace chat with the ask filled in. Nothing runs until you send it.
Get early accessSign in to useMore in Energy and utilities
Reconcile study assumptions across an interconnection queue
For: Interconnection manager at a developer, or a transmission planner reviewing clustered studies
Each project's study rests on a base case, on earlier-queued projects assumed in service, on short-circuit levels and on shared network upgrades.
Check a flexibility contract portfolio against dispatch obligations
For: Portfolio manager at a flexibility aggregator, or a utility demand-response programme lead
Every site contract has its own notice time, duration limit, activation cap, availability window, baseline method and penalty.
Review maintenance and test records across a substation fleet
For: Asset manager or maintenance engineer at a transmission or distribution network operator
Oil analysis reports, breaker timing tests, relay test sheets and battery discharge tests sit inside work orders.