Measured · Study8views
WindTunnel study - cc vs codex
WindTunnel · 50 tasks · 2 harnesses · 2 models
Your setup won the shootout on correctness and efficiency.
Ranked shootout, no baseline arm: cells ranked by the composite score — the benchmark's own graders, then efficiency — claude-code / claude-haiku-4-5-20251001 leads.
Abstract
WindTunnel tested 50 tasks across 2 harness/model cells. Your setup achieved a 95% pass rate, correctly solving the benchmark's tasks. On efficiency, your setup delivered a median cost of $0.08 (mean $0.18) and median token usage of 380.3k (mean 799.4k), completing tasks in a median 4m 29s (mean 5m 44s). The leader across cells is claude-code / claude-haiku-4-5-20251001.
The result
claude-codecodex
Best setup: claude-code / claude-haiku-4-5-20251001 (98.4 of 100)
- 1claude-code / claude-haiku-4-5-2025100198.4best
- 2codex / gpt-5.4-mini81.6
Overall score per cell, 0–100 points (not a pass rate): 75% benchmark pass rate (the benchmark's own graders carry the outcome share — no judge grades or evals on this board) + 25% efficiency (cost · tokens · duration, vs the board's best arm). Absent components renormalize. Best at the top — the board's best setup is tagged.
Leaderboard
Cells ranked by the composite score — outcome (the lenses that graded) × efficiency — absolute values, no baseline arm. Tokens, cost and duration are per-cell medians, the efficiency the score folds.
| # | Harness | Model | Score | Pass rate | Quality | Evals | Goal | Tokens | Cost | Duration | Runs |
|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | claude-code | claude-haiku-4-5-20251001 | 98.4 | 98% | − | − | 100% | 176.5k | $0.03 | 4m 23s | 48/48 |
| 2 | codex | gpt-5.4-mini | 81.6 | 92% | − | − | 100% | 427.3k | $0.13 | 4m 47s | 48/48 |
Pass rate
93.8% (95% CI 88.8–98.8)over 80 graded runs- Task text is withheld from public pages until an admin verifies this licence. Running this benchmark is unaffected.
| Cell | Pass rate | Tasks | Runs | Ran as | Matcher |
|---|---|---|---|---|---|
| claude-code/claude-haiku-4-5-20251001 | 97.5% (95% CI 92.5–100) | 40 | 40 (+8 unmeasured) | with my setup | text_predicate |
| codex/gpt-5.4-mini | 90% (95% CI 80–100) | 40 | 40 (+8 unmeasured) | with my setup | text_predicate |
Each run's answer was extracted from its response and matched against the benchmark's own key — no model in the loop. Two-stage: the mean over repeats per task, then the mean over tasks, with the interval taken over tasks.
Quality × efficiency clusters — normalized per task
Every graded run, standardized WITHIN its task so difficulty cancels out: → right = fewer tokens than the field on the same task, ↑ up = higher quality index (85% the benchmark's own pass/fail + 15% outcome) than the field. The field pools every harness and model, so a setup's left–right position largely reflects its own token habits. Dots are runs; each harness logo is a harness × model × arm centroid — hover it for the model and averages. Up-right wins.
claude-codecodex
Every metric, per harness × model
Benchmark pass rate
- claude-code / claude-haiku-4-5-2025100198%
- codex / gpt-5.4-mini92%
Task completion
- claude-code / claude-haiku-4-5-20251001100%
- codex / gpt-5.4-mini100%
Cost per run
- claude-code / claude-haiku-4-5-20251001$0.03
- codex / gpt-5.4-mini$0.13
Tokens per run
- claude-code / claude-haiku-4-5-20251001176.5k
- codex / gpt-5.4-mini427.3k
Duration per run
- claude-code / claude-haiku-4-5-202510014m 23s
- codex / gpt-5.4-mini4m 47s
Your setup
The publisher has not described this setup — what each group's tools are, who built them, and how the tasks were chosen.
Distributions
Every completed run is one dot — the spread the averages hide. Click a dot to replay that run's journey.
Cost per run
Token composition — average per run
Statistics
| Metric | Arm | n | Mean | Median (pooled) | Min | Max | Std dev |
|---|---|---|---|---|---|---|---|
| Cost | treat | 96/96 | $0.18 | $0.03 | $0.02 | $4.38 | $0.51 |
| Duration | treat | 96/96 | 5m 44s | 4m 36s | 1m 32s | 15m 27s | 3m 12s |
| Total tokens | treat | 96/96 | 799.4k | 179.5k | 26.2k | 18.3M | 2.2M |
| Output tokens | treat | 96/96 | 3.6k | 1k | 392 | 69.8k | 8.2k |
| Cache-read tokens | treat | 96/96 | 750.2k | 172.5k | 10.6k | 17.1M | 2M |
| Cache-write tokens | treat | 96/96 | 3k | 1.7k | 0 | 35.3k | 5k |
| Turns | treat | 96/96 | 10.3 | 9 | 3 | 52 | 7.7 |
n = runs carrying the fact / completed runs in the arm; every statistic runs over present facts only — a missing fact is never counted as 0. Std dev is the sample form (n−1), withheld below n = 2. Every figure here POOLS all runs. The typical-run figures in the abstract, the head-to-head decision and the per-setup panels take each task's median first and then the median across tasks, so a task that ran more often never outweighs one that ran once; the two medians can differ. A promise study's headline (“cut average tokens”) compares per-arm means.
What each setup did
Derived from each run's recorded tool calls — not from a model's description of the run — and aggregated per setup, so a behavior seen across several runs is stated once with its rate. Open a finding to see the runs behind it, each linked to its journey at the step where it happened.
codex · openai · gpt-5.4-mini · Your setup3 findings
- 3 of 48 runs · 3 completed
- 3 of 48 runs · 3 completed
- 1 of 48 runs · 1 completed
Task text withheld
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Tasks
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Runs
Every run is inspectable — open one to replay the agent's journey step by step, with the analyst's read underneath.
Every run, filterable… of 96Show runsHide runs
| Harness | Model | Task | Completed | Pass | Tokens | Cost | Duration |
|---|
Loading runs…
Methodology
What each metric means
- Completed
- The run finished and the harness returned a response. It does NOT mean the answer was correct — a completed run can score zero on quality.
- Denominator: Terminal runs, excluding those killed by our own infrastructure.
- Quality index
- A composite used only in the quality x efficiency plot, built from whichever grader scored this study's runs — the judge's rubric, the pre-registered checks or the benchmark's own verdict — plus whether the run's own outcome was success. The map's caption names the grader and its weights.
- Denominator: Graded runs — runs the study's grader scored.
- Infra-excluded
- A run killed by our own infrastructure. It is a missing measurement, never a loss for the arm, and is excluded from every rate denominator.
- Denominator: Reported as a count beside every affected panel.
- Single arm
- Every task runs once per harness × model cell — a shootout with no baseline arm. The readout is absolute (quality, success, tokens, cost) and the cells are ranked into a leaderboard.
- One run per task
- Every task ran once per cell, in one phrasing. Run-to-run variation is therefore not measured: a single task's difference can be one lucky or unlucky attempt. Every row of the reference table treats the TASKS as the unit: its intervals resample the tasks and its p-values come from a paired test over them, so they describe how much the result depends on which tasks were drawn — not how it would change if the same tasks were run again. With few tasks no difference can reach significance (five tasks cannot go below p = 1/16).
- Sample size
- 2 cells × 50 tasks × 1 arm = 100 runs planned; 96 launched.