Benchmarks,each run with a world model and without.

Anyone can claim their system improves agent decisions. We tested Atlas, Root Cause Analysis and Equinox against public benchmarks that are hard for agents in the same ways real operations are hard: long horizons, incomplete information, and consequences that only show up later.

Incidents repaired
+53pts

47% to 100%, smaller models on live applications.

Root cause, first hit
+11pts

79.4% to 90.6% against the strongest published method.

A month of trading
3.3×

8,709 to 24,076 in net assets, running a shop for thirty days.

Running a business
1 in 3 → none

Runs ending in bankruptcy, and 28% more cash at the end.

Ergodic enhances the agent harness in three ways.

Three of the four run the same agent twice, once with Ergodic components and once without. The root cause evaluation runs differently: there our method is scored against the published methods on that benchmark. What we add sits alongside the agent's own tools: a model of the system it is working in, a way to read back through what happened, and a way to simulate an action before taking it. Which components apply differs by benchmark, and each method note says which were used.

Atlas

The operation held as one connected picture

Atlas keeps every supplier, route, site and order together with what each one depends on and when each thing happened. The connections are part of the record rather than something you rebuild with a query, so an agent can follow an effect from where it started to where it lands, and ask what was known at the time. The same figures sitting in separate tables answer neither question.

RCA

A run back through those events

It lives on Atlas' event history rather than beside it, and walks back in time until it reaches the cause. No second store, no separate copy of the timeline.

Equinox

Scenario generation, built for the call

The part fitted most closely to the goal: it generates the scenarios that matter for that decision, runs them forward and reports what each one is worth.

THE AGENT HARNESSYour agentPlanActObserveTHE SYSTEM IT IS WORKING IN · LIVE CLUSTER, SHOP, BUSINESSWHAT ERGODIC ADDSATLASWhat is true right nowRCAWhy it moved, with the evidenceEQUINOXWhat an action would do nextasksANSWERS ARRIVEAS TOOLS IT CALLSAS TEXT IN ITS PROMPT
Route 01 · tools and skills

The agent calls the framework itself

Atlas, RCA and Equinox are exposed as tools and skills, and the agent decides when to ask. This is how it works in production, and it is the route we use wherever a benchmark allows tool calls.

Route 02 · data in the prompt

The framework's answers are written into the prompt

Some benchmarks forbid extra tool calls, so the framework runs first and its answers are written into the prompt. Less flexible, and it still moves the result: CEO-Bench is won by the arm where the forecast sits inside the decision rather than beside it.

Which components are added, and how they reach the agent, is stated in the method note on each benchmark.

AIOpsLab · Microsoft Research

Finding, naming and repairing live incidents.

A production system breaks. The agent repairs every one it is asked to.

Scenario

On call for a live cluster

A service crashes, a network slows, a configuration goes wrong. Whoever is on call has to find the faulty component, name the failure, and repair it while the system is still serving customers.

Faults are injected into running microservice applications. The agent works through a shell, like a human would.

Measured

Solved, steps taken, compute spent

Three tasks scored separately: find the faulty component, name the failure type, and repair it. A repair only counts if the benchmark re-reads the system afterwards and finds it healthy.

Steps and billed compute are recorded per episode, so speed and cost are visible alongside accuracy.

Result · smaller, cheaper models (Claude Haiku 4.5 and gpt-oss-20b) · three separate evaluations, scored on their own task sets
Start

Something breaks

A fault is injected into a running application. The three tasks below are scored separately, on different task sets.

Live system · real traffic
Step 1

Find it

Which component is at fault?

Alone
59%
Model
76%
34 localisation tasks · mean steps 19.6 → 6.5
Step 2

Name it

What kind of failure is it?

Alone
18%
Model
32%
22 classification tasks · mean steps 17.9 → 8.4
Step 3

Repair it

Is the system healthy again?

Alone
47%
Model
100%
17 repair tasks alone, 15 with the model · scored on cluster state
Result · a near-frontier model (Claude Opus 4.6), same tasks
On its own
1234567891011
14/17found
$5.15compute
With the world model
12
17/17found
$1.14compute

Faults found on the leading model, with the steps and billed compute each arm spent getting there. A leading model rarely needs help reaching the right answer, it needs the evidence gathered: with the current picture already held and the probe results already taken, a long investigation becomes a short confirmation. Steps fall from about eleven to two on these tasks, and compute falls with them.

Setup & Methods · AIOpsLab

Setup

Microsoft AIOpsLab on live Kubernetes, using the DeathStarBench hotel-reservation and social-network testbeds with chaos-injected faults. Baseline is a ReAct agent with shell access; the model arm adds Atlas, RCA and Equinox as queryable layers. Same model, loop, step budget and data in both arms. Models via AWS Bedrock: claude-haiku-4-5 and gpt-oss-20b (smaller), claude-opus-4-6 (leading). Scoring is the benchmark's own: exact match for find and name, cluster state re-read for repair. September 2026.

Numbers and caveats

Repair, smaller models
8/17 → 15/15
Flips for / against
8–0 · McNemar p ≈ 0.008
Find, smaller models
20/34 → 26/34
Name, smaller models
4/22 → 7/22
Leading model, combined
32/42 → 36/42 · 4–0 flips
Network delay and loss
0/8 → 4/4

One seed per model, arm and problem: the repair and leading-model results are lopsided enough to survive noise, the smaller margins need a repeat run on frozen code. Naming a failure requires an exact match against an undocumented label list, which caps every agent in that column. Simulating a repair before applying it costs about 3× the compute on a leading model until prompt caching lands.

RCAEval RE2

Locating the faulty service in a microservice estate.

Two hundred and seventy failures. The right service named first, nine times in ten.

Scenario

Rank the services, faulty one on top

Failures are injected across three microservice systems, with ten minutes of telemetry either side of each injection. The task is to rank the services so the faulty one sits as high as possible.

Six fault types: CPU, memory, disk, socket, network delay and packet loss.

Measured

How often the first answer is right

Accuracy at rank one, at rank three, and averaged over the first five ranks. First-hit accuracy is the one that matters on call: it's the difference between opening the right service and working down a list.

RCA runs back through the events Atlas holds, rather than scoring metrics against each other. On the same cases, that approach ranked the right service first more often than every published method on this benchmark.

Result · held-out test, 180 cases
Best published methodTORAI
79.4%
Atlas + RCAright service ranked first
90.6%
0.932→0.954

Average over the first five ranks, across all 180 held-out cases

0.913→0.983

Online Boutique, the system with the largest gap

Tie

Sock Shop at 0.96 and Train Ticket at 0.92, where the published method is already strong

Setup & Methods · RCAEval RE2

Setup

RCAEval RE2: 270 injected failures across Online Boutique, Sock Shop and Train Ticket, 30 cases per fault type, 60 per system in the held-out split, with ten minutes of telemetry before and after each injection. Baselines run on the same cases: TORAI, N-Sigma, BARO, CIRCA, TraceRCA and MicroRank. Scored with AC@1, AC@3 and Avg@5.

Numbers and caveats

AC@1, held out
79.4% → 90.6%
Avg@5, held out
0.932 → 0.954
Online Boutique
0.913 → 0.983
Sock Shop
0.963 · tie
Train Ticket
0.92 · tie

The method needs distributed tracing and performs worse without call graphs, so it suits estates that already emit traces. Results on the RE3 code-fault split shaped the final algorithm, which means RE2 is the honest read and RE3 is not independent. Some published baselines were not run on all three systems.

MerchantBench

Autonomously running an online store.

An agent runs a shop for a month. It ends with 3.3 times the net assets.

Scenario

A catalogue, fifty shelf slots, thirty days

1,000 products and 50 listing slots. Every day the agent decides what to buy, what to list, what to charge and what to drop, with each commitment closing off options later.

In this arm the agent got the model's map of the shop and nothing else. It set every price itself.

Measured

The money left at the end

Final net assets once in-flight orders settle, plus the operating measures behind it: orders taken, orders delivered, and avoidable failures like stockouts or orders it couldn't pay for.

Paired seeds, so each scenario runs in both arms and wins are counted seed by seed.

Result · net assets at the end of the month, median run
On its own
8,709
With the world model
24,076
3.3×
more net assets
won 11 of 13 head-to-head runs

Equinox found the lever, the agent used it

Running scenarios against this market, Equinox found the exploit in it: price is not the lever, shelf occupancy is. Hand the agent that and it capitalises, keeping more products listed, which it can only do if it knows what is in stock, affordable and arriving in time. Across the runs, money made tracked live listings far more closely than price.

383→679

Orders per month

59%→74%

Orders delivered

0.27→0.10

Avoidable failure rate

Setup & Methods · MerchantBench

Setup

MerchantBench (arXiv 2607.28956), 30-day scenario: 720 simulated hours, 1,000 SKUs, 50 listing slots, gpt-oss-120b via AWS Bedrock. Both arms use the same agent loop, step budget and read batch; the model arm adds a queryable map of the shop. Score is final net assets read after in-flight orders settle. Paired seeds with exact Wilcoxon tests; ratios are geometric means and amounts are medians.

Numbers and caveats

Map only vs same agent
11 of 13 · 3.30× · p = 0.002
How 3.3× is derived
Geometric mean of the paired per-seed ratios; amounts quoted are medians
Map only vs stock baseline
5 of 6 · 2.30× · p = 0.063
Listings correlation
+0.83 listings · +0.20 markup
Diagnosis layer added
1 of 4 · no gain yet

With full planning and pricing the shop made 15 times more, but the scenario puts no ceiling on price, so that figure is inflated and we don't count it. Two scoring bugs that favoured us were found and fixed, and all 123 runs rescored. Adding the diagnosis layer on top of the map has not paid off yet in this scenario.

CEO-Bench

Running a simulated software company for a quarter.

The model alone runs the company into the ground once in every three attempts. With a world model underneath it, never.

Scenario

A simulated SaaS business, run to the end

An agent runs a software company for 91 simulated days, setting price, hiring, R&D spend and advertising budget, against customer segments whose behaviour it can only learn by acting.

Three arms: no tools, tools offered as advice, and the forecast placed in the decision itself.

Measured

Cash in the bank on the last day

Final cash after 91 days, across 20 held-out seeds run three times each, 180 runs in total, so agent sampling noise can be separated from the effect of the tools.

Repeated-measures paired t-tests, within-seed spread of $80k to $195k.

Result · mean cash on day 91, 180 runs
No tools
$287,615
Tools offered
$338,759
Forecast in the decision
$367,198
1 in 3→none

Runs where the business went bankrupt before day 91

+28%

More cash at the end of the run, +$79.6k against the baseline, p = 0.037

p ≈ 0.20

When the tools are offered as advice, the gain is inside sampling noise

The agent had all three parts: Atlas for exact recall of what it had already done and what each segment did next, Equinox for the cash forecast under each option, RCA for reading back through the decisions that went badly. It priced and spent better for it, and held on to more of its customers. The finding worth carrying is about delivery: handing an agent a forecast it may consult moved the mean but sat inside sampling noise, while putting that forecast where the decision is made produced the measurable gain. The comparison between the two deliveries is not yet resolved, and the stronger arm is a diagnostic ceiling rather than a leaderboard-valid entry.

Setup & Methods · CEO-Bench

Setup

Claude Haiku 4.5 at temperature 1.0 operating the benchmark's simulated business for 91 days. Components: Atlas for exact recall of what had already happened, RCA as Dirichlet-multinomial churn posteriors by segment, Equinox as a Monte Carlo cash forecaster with 95% intervals at four horizons. 20 held-out seeds, three development seeds excluded, each seed run three times per arm.

Numbers and caveats

No tools
$287,615
Tools offered
$338,759 · p ≈ 0.20
Forecast in the decision
$367,198 · p = 0.037
Paired t, forced vs baseline
t = 2.24

The forced arm is a diagnostic ceiling rather than a leaderboard-valid submission, and the mandatory-review design that would be valid was built but not yet run at scale. Tools-offered against baseline, and forced against tools-offered, remain unresolved and need several times the seeds. Three correctness bugs were found and fixed during the evaluation: price reconstruction, forecast interval coverage (0.024 to 0.944), and burn-rate staleness. Arms carrying tools showed higher sampling variance than the baseline, so advisory access also makes behaviour less consistent.

Run this on your own operation.

Bring the decision your team argues about most. We'll build the model around it and simulate it forward with your own numbers underneath.

Talk To An Expert