
DAI Bench
AI coding agents, run through each vendor's own command line tool on a paid subscription and scored by tests.
Standings
| Engine | Plan | CLI version | Status |
|---|---|---|---|
| Claude Code | Claude Max | 2.1.288 | Harness checked 2026-10-03, awaiting first run |
| Codex CLI | ChatGPT | 0.160.0 | Harness checked 2026-10-03, awaiting first run |
| Gemini CLI | Google account | not recorded | Not checked yet |
| Ollama | none (local GPU) | not recorded | Not checked yet |
One entry, three parts
Every result here belongs to an entry: one engine, one model and one effort level. The same model at a higher effort is a different entry, because that is a different thing to pay for and wait for.
A release run covers every effort level the engine offers. The weekly Drift check runs two of them per model: the engine's default and its highest.
- Claude Codemodellow
- Claude Codemodelmedium
- Claude Codemodelhighdefault
- Claude Codemodelxhigh
- Claude Codemodelmax
- Codex CLI
- low, medium, high, xhigh, max, ultra (default medium)
- Gemini CLI
- no effort setting, one entry per model
- Ollama
- false, true, low, medium, high (default true)
What is measured
Five axes, each scored 0 to 100. The overall score is their weighted mean, with the weights published here and in weights.json.
| Axis | What it measures | Scored by | Weight |
|---|---|---|---|
| Coding | Functions and bug fixes in Rust, TypeScript and Python. The agent edits a fixture. | Hidden unit tests | 35 |
| Agentic | Multi-step tasks across several files, with a written spec and a mock JSON API served locally. | Final-state check | 25 |
| Fidelity | Strict output files: JSON schemas, length limits, negative constraints, format rules. | Validators | 15 |
| Long context | Finding and combining facts across a folder of docs at 32k and 128k tokens. | Exact or set match | 15 |
| Speed and usage | Wall-clock time per task, and the tokens and cost each CLI reports. | Measured | 10 |
| Overall | Weighted mean of the five axes | 100 |
Every score comes from a program that inspects the task folder after the agent exits: tests, validators, file state. No model judges another. Each axis carries a 95% confidence interval from repeated runs and item resampling.
Task items stay private so they cannot leak into training data. Each run publishes the hash of every item and a public sample.
Drift
Has an entry got worse since release? Each tracked entry is checked every week against its own first runs.
- Baseline
- Five canary runs in the first seven days after an entry is added. The control band is the mean plus or minus two standard deviations, never narrower than two points either side.
- Weekly check
- The same 40 canary items, compared item by item with the baseline: an exact McNemar test for pass or fail items, a bootstrap interval for scores, and a Holm correction across entries.
- Harness change
- A new CLI version is drawn as a marker. The next check is labelled harness changed and does not count toward a verdict, so a worse CLI update is told apart from a worse model.
How it is run
Each engine is the vendor's own command line tool, run headless as its documentation describes, on the plan named in the standings. The runner drops API key variables from the tool's environment, so a run can only use the subscription login.
Every task starts in a fresh folder that holds only that task's files. The agent works there under a wall-clock limit. On a timeout the runner stops its own process tree by process id, and nothing else.
The tools offer no temperature or seed setting, so runs are repeated instead and every score carries its measured spread. Each run records the tool's version, the model flag, the plan and the usage the tool reports.
A run stops at a tenth of the plan's weekly allowance. Invocations follow each vendor's documentation: Claude Code, Codex CLI, Gemini CLI and Ollama.
Check it yourself
As this chapter came into view, the page downloaded every data file it was built from, hashed each one with BLAKE2b-512 here in the browser, and checked the minisign signature over the list.
- data/engines.json3,705 bytes5cdb34673ec683ad…ad984dc5waiting
- data/weights.json95 bytesb159fa56caa38414…2381e1e5waiting
- data/brands.json15,633 bytese3dca3f64df89fa6…8dbab8dfwaiting
- data/models.json3 bytes16df9553704b6efe…be40cf30waiting
- manifest.jsonsignaturekey AD2261F640A15481waiting
Checking
Colophon
Check it yourself
Download manifest.json and manifest.json.minisig, then run:
minisign -Vm manifest.json -P RWSBVKFA9mEirbS11TtAIbOBMsh3L1aXT8debpA4Cyi4NgQ8o7EyoGA9The manifest lists each data file with its size and BLAKE2b-512 hash. b2sum prints the same hash.
Public key
RWSBVKFA9mEirbS11TtAIbOBMsh3L1aXT8debpA4Cyi4NgQ8o7EyoGA9
Key id AD2261F640A15481. Also at minisign.pub.
DAI Bench
Built 2026-10-03 15:53 UTC.
Set in IBM Plex Sans and IBM Plex Mono. A static site: no server code, no database, no cookies.