Example data · mock provider · virtual time

A comparison you can reproduce

Sanitized Bench Card · CC-Reconcile · published sample

One agent vs adaptive team

StrategyVerifiedVirtual timeCost
One strong agent96.7%2668 ms$0.0035
Adaptive team, at most 4 agents93.3%1431 ms$0.0062

30 cases · time and cost are medians · success difference, team minus one agent, −0.033 · 95% interval −0.100 to 0.000 · resampling unit: case (paired, 2000 resamples).

What the sample shows

In this mock sample the team finished sooner in virtual time and did not verify more cases than one agent. Over the 28 cases where both succeeded, the team was a median 1.81x sooner. Per run it cost $0.0062 against $0.0035(1.74x): both far below a cent, so cost does not decide this comparison. The constants behind those numbers are arbitrary simulator settings, so they describe mechanics, not models.

What this card does and does not show

Only seeded mock results and virtual time. Real-model performance is not established. Private source text is excluded.

  • MOCK PROVIDER. Latency, cost and mistakes come from a seeded simulator with arbitrary constants. Times are virtual milliseconds.
  • This measures orchestration mechanics only. It says nothing about real model quality, real latency or real prices.
  • Aggregation is a deterministic merge with no model call, so its provider cost is zero by construction.
  • Only B1 and B5 are implemented. B0, B2, B3, B4 and B6 to B8 are not, so no claim of advantage over them exists.

Reproducibility

Suite
CC-Reconcile
Dataset version
cc-reconcile-1
Sample size
30 cases x 2 strategies
Execution
mock provider, virtual time
Resource envelope
at most 4 agents, $0.2000 authorized per run
Stop rule
fixed: 30 cases x 2 strategies, no early stopping
Engine generation
gen_seed_aa981c6c05f9
Generated
2026-09-20T10:29:39.257Z
Source availability
Generated fixture. No private sources.
Live model profile
not supplied
Reproducibility identifier
sha256 8745c7774472cccf95dcc5d5ba3636c87fabf7cd677bcb6e781563d06fb3144b

Reproduce locally, no credentials needed:

npm run ccbench -- compare --cases 30