A comparison you can reproduce
Sanitized Bench Card · CC-Reconcile · published sample
One agent vs adaptive team
| Strategy | Verified | Virtual time | Cost |
|---|---|---|---|
| One strong agent | 96.7% | 2668 ms | $0.0035 |
| Adaptive team, at most 4 agents | 93.3% | 1431 ms | $0.0062 |
30 cases · time and cost are medians · success difference, team minus one agent, −0.033 · 95% interval −0.100 to 0.000 · resampling unit: case (paired, 2000 resamples).
What the sample shows
In this mock sample the team finished sooner in virtual time and did not verify more cases than one agent. Over the 28 cases where both succeeded, the team was a median 1.81x sooner. Per run it cost $0.0062 against $0.0035(1.74x): both far below a cent, so cost does not decide this comparison. The constants behind those numbers are arbitrary simulator settings, so they describe mechanics, not models.
What this card does and does not show
Only seeded mock results and virtual time. Real-model performance is not established. Private source text is excluded.
- MOCK PROVIDER. Latency, cost and mistakes come from a seeded simulator with arbitrary constants. Times are virtual milliseconds.
- This measures orchestration mechanics only. It says nothing about real model quality, real latency or real prices.
- Aggregation is a deterministic merge with no model call, so its provider cost is zero by construction.
- Only B1 and B5 are implemented. B0, B2, B3, B4 and B6 to B8 are not, so no claim of advantage over them exists.
Reproducibility
- Suite
- CC-Reconcile
- Dataset version
- cc-reconcile-1
- Sample size
- 30 cases x 2 strategies
- Execution
- mock provider, virtual time
- Resource envelope
- at most 4 agents, $0.2000 authorized per run
- Stop rule
- fixed: 30 cases x 2 strategies, no early stopping
- Engine generation
- gen_seed_aa981c6c05f9
- Generated
- 2026-09-20T10:29:39.257Z
- Source availability
- Generated fixture. No private sources.
- Live model profile
- not supplied
- Reproducibility identifier
- sha256 8745c7774472cccf95dcc5d5ba3636c87fabf7cd677bcb6e781563d06fb3144b
Reproduce locally, no credentials needed:
npm run ccbench -- compare --cases 30