DAI Bench

AI coding agents, run through each vendor's own command line tool on a paid subscription and scored by tests.

Standings

No scored runs yet. Scores appear here when the first signed run is published.
EnginePlanCLI versionStatus
Claude CodeClaude Max2.1.288Harness checked 2026-10-03, awaiting first run
Codex CLIChatGPT0.160.0Harness checked 2026-10-03, awaiting first run
Gemini CLIGoogle accountnot recordedNot checked yet
Ollamanone (local GPU)not recordedNot checked yet

One entry, three parts

Every result here belongs to an entry: one engine, one model and one effort level. The same model at a higher effort is a different entry, because that is a different thing to pay for and wait for.

A release run covers every effort level the engine offers. The weekly Drift check runs two of them per model: the engine's default and its highest.

The effort ladder for Claude Code, before any model is added
  1. Claude Codemodellow
  2. Claude Codemodelmedium
  3. Claude Codemodelhighdefault
  4. Claude Codemodelxhigh
  5. Claude Codemodelmax
Codex CLI
low, medium, high, xhigh, max, ultra (default medium)
Gemini CLI
no effort setting, one entry per model
Ollama
false, true, low, medium, high (default true)

What is measured

Five axes, each scored 0 to 100. The overall score is their weighted mean, with the weights published here and in weights.json.

Axes, how each is scored, and its weight in the overall score
AxisWhat it measuresScored byWeight
CodingFunctions and bug fixes in Rust, TypeScript and Python. The agent edits a fixture.Hidden unit tests35
AgenticMulti-step tasks across several files, with a written spec and a mock JSON API served locally.Final-state check25
FidelityStrict output files: JSON schemas, length limits, negative constraints, format rules.Validators15
Long contextFinding and combining facts across a folder of docs at 32k and 128k tokens.Exact or set match15
Speed and usageWall-clock time per task, and the tokens and cost each CLI reports.Measured10
OverallWeighted mean of the five axes100

Every score comes from a program that inspects the task folder after the agent exits: tests, validators, file state. No model judges another. Each axis carries a 95% confidence interval from repeated runs and item resampling.

Task items stay private so they cannot leak into training data. Each run publishes the hash of every item and a public sample.

Drift

Has an entry got worse since release? Each tracked entry is checked every week against its own first runs.

Baseline, 5 runsCLI updatedDegraded
Diagram of the rule, not a measurement. No entry has a baseline yet. Shaded: the control band. Dashed: a CLI update. The last two checks fall below the band, which reads Degraded.
Baseline
Five canary runs in the first seven days after an entry is added. The control band is the mean plus or minus two standard deviations, never narrower than two points either side.
Weekly check
The same 40 canary items, compared item by item with the baseline: an exact McNemar test for pass or fail items, a bootstrap interval for scores, and a Holm correction across entries.
Harness change
A new CLI version is drawn as a marker. The next check is labelled harness changed and does not count toward a verdict, so a worse CLI update is told apart from a worse model.

How it is run

Each engine is the vendor's own command line tool, run headless as its documentation describes, on the plan named in the standings. The runner drops API key variables from the tool's environment, so a run can only use the subscription login.

Every task starts in a fresh folder that holds only that task's files. The agent works there under a wall-clock limit. On a timeout the runner stops its own process tree by process id, and nothing else.

The tools offer no temperature or seed setting, so runs are repeated instead and every score carries its measured spread. Each run records the tool's version, the model flag, the plan and the usage the tool reports.

A run stops at a tenth of the plan's weekly allowance. Invocations follow each vendor's documentation: Claude Code, Codex CLI, Gemini CLI and Ollama.

Check it yourself

As this chapter came into view, the page downloaded every data file it was built from, hashed each one with BLAKE2b-512 here in the browser, and checked the minisign signature over the list.

  1. data/engines.json3,705 bytes5cdb34673ec683ad…ad984dc5waiting
  2. data/weights.json95 bytesb159fa56caa38414…2381e1e5waiting
  3. data/brands.json15,633 bytese3dca3f64df89fa6…8dbab8dfwaiting
  4. data/models.json3 bytes16df9553704b6efe…be40cf30waiting
  5. manifest.jsonsignaturekey AD2261F640A15481waiting

Checking

Colophon

Check it yourself

Download manifest.json and manifest.json.minisig, then run:

minisign -Vm manifest.json -P RWSBVKFA9mEirbS11TtAIbOBMsh3L1aXT8debpA4Cyi4NgQ8o7EyoGA9

The manifest lists each data file with its size and BLAKE2b-512 hash. b2sum prints the same hash.

Public key

RWSBVKFA9mEirbS11TtAIbOBMsh3L1aXT8debpA4Cyi4NgQ8o7EyoGA9

Key id AD2261F640A15481. Also at minisign.pub.

Data files

DAI Bench

Built 2026-10-03 15:53 UTC.

Set in IBM Plex Sans and IBM Plex Mono. A static site: no server code, no database, no cookies.