Delivered · Study4views
WindTunnel A/B — Playwright vs WebMCP (50 tasks)
WebMCP vs Playwright · 50 tasks · 2 harnesses · 2 models
WebMCP matches Playwright's correctness while cutting cost, tokens, and time across the board.
Both arms reached the same result, and the treatment got there with 41% lower cost and 40% fewer tokens.
Abstract
Both WebMCP and Playwright achieved a 99% pass rate on the 50-task benchmark, with no difference in correctness. WebMCP delivered substantial efficiency gains: a median cost of $0.02 (mean $0.03) against Playwright at $0.04 (mean $0.07), representing -41.5% cost reduction. Token usage came in at a median 118.4k (mean 129.7k) for WebMCP versus 196.1k (mean 373.9k) for Playwright, a -39.7% reduction. Execution duration was a median 33.5s (mean 36.7s) for WebMCP against 44.7s (mean 59.6s) for Playwright, a -25.1% improvement. These efficiency gains held across both harness and model cells, though harness and model are not independently crossed in this design.
The result
PlaywrightWebMCP
Best arm overall: WebMCP (94.9 vs 87.7 of 100)
Same outcome on both arms — the whole gap is efficiency (cost · tokens · duration).
Best setup: codex / gpt-5.4-mini · WebMCP (98.4 of 100)
- 1codex / gpt-5.4-miniWebMCP98.4best
- 2claude-code / claude-haiku-4-5-20251001WebMCP93.6
- 3codex / gpt-5.4-miniPlaywright92.0
- 4claude-code / claude-haiku-4-5-20251001Playwright86.8
Overall score per arm, 0–100 points (not a pass rate): 75% benchmark pass rate (the benchmark's own graders carry the outcome share — no judge grades or evals on this board) + 25% efficiency (cost · tokens · duration, vs the board's best arm). Absent components renormalize. Best at the top — the board's best setup is tagged.
Quality × efficiency clusters — normalized per task
Every graded run, standardized WITHIN its task so difficulty cancels out: → right = fewer tokens than the field on the same task, ↑ up = higher quality index (85% the benchmark's own pass/fail + 15% outcome) than the field. The field pools every harness and model, so a setup's left–right position largely reflects its own token habits. Dots are runs; each harness logo is a harness × model × arm centroid — hover it for the model and averages. Up-right wins.
4 runs not plotted (missing a grade or token count — infra failures included).
Every metric, per harness × model
Benchmark pass rate
- claude-code / claude-haiku-4-5-20251001100% → 100%even
- codex / gpt-5.4-mini98% → 98%even
Task completion
- claude-code / claude-haiku-4-5-20251001100% → 100%even
- codex / gpt-5.4-mini100% → 100%even
Cost per run
- codex / gpt-5.4-mini$0.03 → $0.0230% better
- claude-code / claude-haiku-4-5-20251001$0.05 → $0.0343% better
Tokens per run
- codex / gpt-5.4-mini96.5k → 71k26% better
- claude-code / claude-haiku-4-5-20251001297.2k → 177.1k40% better
Duration per run
- claude-code / claude-haiku-4-5-2025100146s → 32.6s29% better
- codex / gpt-5.4-mini41.6s → 32.7s21% better
Playwright
WebMCP
The publisher has not described this setup — what each group's tools are, who built them, and how the tasks were chosen.
Distributions
Every completed run is one dot — the spread the averages hide. Click a dot to replay that run's journey.
Cost per run
Token composition — average per run
Run outcomes
How each run ended — whether the agent finished, not whether its answer passed. Whether answers passed is in the results above.
Statistics
| Metric | Arm | n | Mean | Median (pooled) | Min | Max | Std dev |
|---|---|---|---|---|---|---|---|
| Cost | base | 100/100 | $0.07 | $0.04 | $0.01 | $0.63 | $0.10 |
| treat | 100/100 | $0.03 | $0.02 | $0.01 | $0.07 | $0.01 | |
| Duration | base | 100/100 | 59.6s | 42.6s | 24.4s | 4m 18s | 43.5s |
| treat | 100/100 | 36.7s | 32.7s | 26.8s | 2m 00s | 12s | |
| Total tokens | base | 100/100 | 373.9k | 166.9k | 34.1k | 3.1M | 549.6k |
| treat | 100/100 | 129.7k | 107.5k | 58k | 509.7k | 74.1k | |
| Output tokens | base | 100/100 | 1.7k | 929 | 188 | 15.6k | 2.2k |
| treat | 100/100 | 761 | 645 | 350 | 2.9k | 438 | |
| Cache-read tokens | base | 100/100 | 350.6k | 148k | 21.8k | 3M | 527.4k |
| treat | 100/100 | 121.3k | 101k | 32.9k | 500k | 75.1k | |
| Cache-write tokens | base | 100/100 | 8.9k | 2.2k | 0 | 81.6k | 14.9k |
| treat | 100/100 | 2.9k | 1.7k | 0 | 36.7k | 4.4k | |
| Turns | base | 100/100 | 15.2 | 5.5 | 2 | 98 | 19.8 |
| treat | 100/100 | 8.3 | 7 | 3 | 29 | 5.4 |
n = runs carrying the fact / completed runs in the arm; every statistic runs over present facts only — a missing fact is never counted as 0. Std dev is the sample form (n−1), withheld below n = 2. Every figure here POOLS all runs. The typical-run figures in the abstract, the head-to-head decision and the per-setup panels take each task's median first and then the median across tasks, so a task that ran more often never outweighs one that ran once; the two medians can differ. A promise study's headline (“cut average tokens”) compares per-arm means.
Task text withheld
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Tasks
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Runs
Every run is inspectable — open one to replay the agent's journey step by step, with the analyst's read underneath.
Every run, filterable… of 200Show runsHide runs
| Harness | Model | Task | Arm | Completed | Pass | Tokens | Cost | Duration |
|---|
Loading runs…
Methodology
What each metric means
- Completed
- The run finished and the harness returned a response. It does NOT mean the answer was correct — a completed run can score zero on quality.
- Denominator: Terminal runs, excluding those killed by our own infrastructure.
- Quality index
- A composite used only in the quality x efficiency plot, built from whichever grader scored this study's runs — the judge's rubric, the pre-registered checks or the benchmark's own verdict — plus whether the run's own outcome was success. The map's caption names the grader and its weights.
- Denominator: Graded runs — runs the study's grader scored.
- Infra-excluded
- A run killed by our own infrastructure. It is a missing measurement, never a loss for the arm, and is excluded from every rate denominator.
- Denominator: Reported as a count beside every affected panel.
- A/B arms
- Every task runs twice per harness × model cell — once as Playwright, once as WebMCP — on the same prompt, same model, cold start for both arms. Both arms are real configurations.
- One run per task
- Every task ran once per cell and arm, in one phrasing. Run-to-run variation is therefore not measured: a single task's difference can be one lucky or unlucky attempt. Every row of the reference table treats the TASKS as the unit: its intervals resample the tasks and its p-values come from a paired test over them, so they describe how much the result depends on which tasks were drawn — not how it would change if the same tasks were run again. With few tasks no difference can reach significance (five tasks cannot go below p = 1/16).
- Sample size
- 2 cells × 50 tasks × 2 arms = 200 runs.