Benchmarks,each run with a world model and without.
Anyone can claim their system improves agent decisions. We tested Atlas, Root Cause Analysis and Equinox against public benchmarks that are hard for agents in the same ways real operations are hard: long horizons, incomplete information, and consequences that only show up later.
Ergodic enhances the agent harness in three ways.
Three of the four run the same agent twice, once with Ergodic components and once without. The root cause evaluation runs differently: there our method is scored against the published methods on that benchmark. What we add sits alongside the agent's own tools: a model of the system it is working in, a way to read back through what happened, and a way to simulate an action before taking it. Which components apply differs by benchmark, and each method note says which were used.
The operation held as one connected picture
Atlas keeps every supplier, route, site and order together with what each one depends on and when each thing happened. The connections are part of the record rather than something you rebuild with a query, so an agent can follow an effect from where it started to where it lands, and ask what was known at the time. The same figures sitting in separate tables answer neither question.
A run back through those events
It lives on Atlas' event history rather than beside it, and walks back in time until it reaches the cause. No second store, no separate copy of the timeline.
Scenario generation, built for the call
The part fitted most closely to the goal: it generates the scenarios that matter for that decision, runs them forward and reports what each one is worth.
The agent calls the framework itself
Atlas, RCA and Equinox are exposed as tools and skills, and the agent decides when to ask. This is how it works in production, and it is the route we use wherever a benchmark allows tool calls.
The framework's answers are written into the prompt
Some benchmarks forbid extra tool calls, so the framework runs first and its answers are written into the prompt. Less flexible, and it still moves the result: CEO-Bench is won by the arm where the forecast sits inside the decision rather than beside it.
Which components are added, and how they reach the agent, is stated in the method note on each benchmark.
Finding, naming and repairing live incidents.
A production system breaks. The agent repairs every one it is asked to.
On call for a live cluster
A service crashes, a network slows, a configuration goes wrong. Whoever is on call has to find the faulty component, name the failure, and repair it while the system is still serving customers.
Faults are injected into running microservice applications. The agent works through a shell, like a human would.
Solved, steps taken, compute spent
Three tasks scored separately: find the faulty component, name the failure type, and repair it. A repair only counts if the benchmark re-reads the system afterwards and finds it healthy.
Steps and billed compute are recorded per episode, so speed and cost are visible alongside accuracy.
Something breaks
A fault is injected into a running application. The three tasks below are scored separately, on different task sets.
Find it
Which component is at fault?
Name it
What kind of failure is it?
Repair it
Is the system healthy again?
Faults found on the leading model, with the steps and billed compute each arm spent getting there. A leading model rarely needs help reaching the right answer, it needs the evidence gathered: with the current picture already held and the probe results already taken, a long investigation becomes a short confirmation. Steps fall from about eleven to two on these tasks, and compute falls with them.
Setup & Methods · AIOpsLab
Setup
Microsoft AIOpsLab on live Kubernetes, using the DeathStarBench hotel-reservation and social-network testbeds with chaos-injected faults. Baseline is a ReAct agent with shell access; the model arm adds Atlas, RCA and Equinox as queryable layers. Same model, loop, step budget and data in both arms. Models via AWS Bedrock: claude-haiku-4-5 and gpt-oss-20b (smaller), claude-opus-4-6 (leading). Scoring is the benchmark's own: exact match for find and name, cluster state re-read for repair. September 2026.
Numbers and caveats
- Repair, smaller models
- 8/17 → 15/15
- Flips for / against
- 8–0 · McNemar p ≈ 0.008
- Find, smaller models
- 20/34 → 26/34
- Name, smaller models
- 4/22 → 7/22
- Leading model, combined
- 32/42 → 36/42 · 4–0 flips
- Network delay and loss
- 0/8 → 4/4
One seed per model, arm and problem: the repair and leading-model results are lopsided enough to survive noise, the smaller margins need a repeat run on frozen code. Naming a failure requires an exact match against an undocumented label list, which caps every agent in that column. Simulating a repair before applying it costs about 3× the compute on a leading model until prompt caching lands.
Locating the faulty service in a microservice estate.
Two hundred and seventy failures. The right service named first, nine times in ten.
Rank the services, faulty one on top
Failures are injected across three microservice systems, with ten minutes of telemetry either side of each injection. The task is to rank the services so the faulty one sits as high as possible.
Six fault types: CPU, memory, disk, socket, network delay and packet loss.
How often the first answer is right
Accuracy at rank one, at rank three, and averaged over the first five ranks. First-hit accuracy is the one that matters on call: it's the difference between opening the right service and working down a list.
RCA runs back through the events Atlas holds, rather than scoring metrics against each other. On the same cases, that approach ranked the right service first more often than every published method on this benchmark.
Average over the first five ranks, across all 180 held-out cases
Online Boutique, the system with the largest gap
Sock Shop at 0.96 and Train Ticket at 0.92, where the published method is already strong
Setup & Methods · RCAEval RE2
Setup
RCAEval RE2: 270 injected failures across Online Boutique, Sock Shop and Train Ticket, 30 cases per fault type, 60 per system in the held-out split, with ten minutes of telemetry before and after each injection. Baselines run on the same cases: TORAI, N-Sigma, BARO, CIRCA, TraceRCA and MicroRank. Scored with AC@1, AC@3 and Avg@5.
Numbers and caveats
- AC@1, held out
- 79.4% → 90.6%
- Avg@5, held out
- 0.932 → 0.954
- Online Boutique
- 0.913 → 0.983
- Sock Shop
- 0.963 · tie
- Train Ticket
- 0.92 · tie
The method needs distributed tracing and performs worse without call graphs, so it suits estates that already emit traces. Results on the RE3 code-fault split shaped the final algorithm, which means RE2 is the honest read and RE3 is not independent. Some published baselines were not run on all three systems.
Autonomously running an online store.
An agent runs a shop for a month. It ends with 3.3 times the net assets.
A catalogue, fifty shelf slots, thirty days
1,000 products and 50 listing slots. Every day the agent decides what to buy, what to list, what to charge and what to drop, with each commitment closing off options later.
In this arm the agent got the model's map of the shop and nothing else. It set every price itself.
The money left at the end
Final net assets once in-flight orders settle, plus the operating measures behind it: orders taken, orders delivered, and avoidable failures like stockouts or orders it couldn't pay for.
Paired seeds, so each scenario runs in both arms and wins are counted seed by seed.
Equinox found the lever, the agent used it
Running scenarios against this market, Equinox found the exploit in it: price is not the lever, shelf occupancy is. Hand the agent that and it capitalises, keeping more products listed, which it can only do if it knows what is in stock, affordable and arriving in time. Across the runs, money made tracked live listings far more closely than price.
Orders per month
Orders delivered
Avoidable failure rate
Setup & Methods · MerchantBench
Setup
MerchantBench (arXiv 2607.28956), 30-day scenario: 720 simulated hours, 1,000 SKUs, 50 listing slots, gpt-oss-120b via AWS Bedrock. Both arms use the same agent loop, step budget and read batch; the model arm adds a queryable map of the shop. Score is final net assets read after in-flight orders settle. Paired seeds with exact Wilcoxon tests; ratios are geometric means and amounts are medians.
Numbers and caveats
- Map only vs same agent
- 11 of 13 · 3.30× · p = 0.002
- How 3.3× is derived
- Geometric mean of the paired per-seed ratios; amounts quoted are medians
- Map only vs stock baseline
- 5 of 6 · 2.30× · p = 0.063
- Listings correlation
- +0.83 listings · +0.20 markup
- Diagnosis layer added
- 1 of 4 · no gain yet
With full planning and pricing the shop made 15 times more, but the scenario puts no ceiling on price, so that figure is inflated and we don't count it. Two scoring bugs that favoured us were found and fixed, and all 123 runs rescored. Adding the diagnosis layer on top of the map has not paid off yet in this scenario.
Running a simulated software company for a quarter.
The model alone runs the company into the ground once in every three attempts. With a world model underneath it, never.
A simulated SaaS business, run to the end
An agent runs a software company for 91 simulated days, setting price, hiring, R&D spend and advertising budget, against customer segments whose behaviour it can only learn by acting.
Three arms: no tools, tools offered as advice, and the forecast placed in the decision itself.
Cash in the bank on the last day
Final cash after 91 days, across 20 held-out seeds run three times each, 180 runs in total, so agent sampling noise can be separated from the effect of the tools.
Repeated-measures paired t-tests, within-seed spread of $80k to $195k.
Runs where the business went bankrupt before day 91
More cash at the end of the run, +$79.6k against the baseline, p = 0.037
When the tools are offered as advice, the gain is inside sampling noise
The agent had all three parts: Atlas for exact recall of what it had already done and what each segment did next, Equinox for the cash forecast under each option, RCA for reading back through the decisions that went badly. It priced and spent better for it, and held on to more of its customers. The finding worth carrying is about delivery: handing an agent a forecast it may consult moved the mean but sat inside sampling noise, while putting that forecast where the decision is made produced the measurable gain. The comparison between the two deliveries is not yet resolved, and the stronger arm is a diagnostic ceiling rather than a leaderboard-valid entry.
Setup & Methods · CEO-Bench
Setup
Claude Haiku 4.5 at temperature 1.0 operating the benchmark's simulated business for 91 days. Components: Atlas for exact recall of what had already happened, RCA as Dirichlet-multinomial churn posteriors by segment, Equinox as a Monte Carlo cash forecaster with 95% intervals at four horizons. 20 held-out seeds, three development seeds excluded, each seed run three times per arm.
Numbers and caveats
- No tools
- $287,615
- Tools offered
- $338,759 · p ≈ 0.20
- Forecast in the decision
- $367,198 · p = 0.037
- Paired t, forced vs baseline
- t = 2.24
The forced arm is a diagnostic ceiling rather than a leaderboard-valid submission, and the mandatory-review design that would be valid was built but not yet run at scale. Tools-offered against baseline, and forced against tools-offered, remain unresolved and need several times the seeds. Three correctness bugs were found and fixed during the evaluation: price reconstruction, forecast interval coverage (0.024 to 0.944), and burn-rate staleness. Arms carrying tools showed higher sampling variance than the baseline, so advisory access also makes behaviour less consistent.
Run this on your own operation.
Bring the decision your team argues about most. We'll build the model around it and simulate it forward with your own numbers underneath.
Talk To An Expert