Delivered · Study4views

WindTunnel A/B — Playwright vs WebMCP (50 tasks)

WebMCP vs Playwright · 50 tasks · 2 harnesses · 2 models

WebMCP matches Playwright's correctness while cutting cost, tokens, and time across the board.

Both arms reached the same result, and the treatment got there with 41% lower cost and 40% fewer tokens.

Abstract

Both WebMCP and Playwright achieved a 99% pass rate on the 50-task benchmark, with no difference in correctness. WebMCP delivered substantial efficiency gains: a median cost of $0.02 (mean $0.03) against Playwright at $0.04 (mean $0.07), representing -41.5% cost reduction. Token usage came in at a median 118.4k (mean 129.7k) for WebMCP versus 196.1k (mean 373.9k) for Playwright, a -39.7% reduction. Execution duration was a median 33.5s (mean 36.7s) for WebMCP against 44.7s (mean 59.6s) for Playwright, a -25.1% improvement. These efficiency gains held across both harness and model cells, though harness and model are not independently crossed in this design.

The result

PlaywrightWebMCP

Best arm overall: WebMCP (94.9 vs 87.7 of 100)

Same outcome on both arms — the whole gap is efficiency (cost · tokens · duration).

Best setup: codex / gpt-5.4-mini · WebMCP (98.4 of 100)

  1. 1codex / gpt-5.4-miniWebMCP98.4best
  2. 2claude-code / claude-haiku-4-5-20251001WebMCP93.6
  3. 3codex / gpt-5.4-miniPlaywright92.0
  4. 4claude-code / claude-haiku-4-5-20251001Playwright86.8

Overall score per arm, 0–100 points (not a pass rate): 75% benchmark pass rate (the benchmark's own graders carry the outcome share — no judge grades or evals on this board) + 25% efficiency (cost · tokens · duration, vs the board's best arm). Absent components renormalize. Best at the top — the board's best setup is tagged.

Quality × efficiency clusters — normalized per task

Every graded run, standardized WITHIN its task so difficulty cancels out: → right = fewer tokens than the field on the same task, ↑ up = higher quality index (85% the benchmark's own pass/fail + 15% outcome) than the field. The field pools every harness and model, so a setup's left–right position largely reflects its own token habits. Dots are runs; each harness logo is a harness × model × arm centroid — hover it for the model and averages. Up-right wins.

Label
Task
Arm
Harness
Model
196 of 196 runs match
Show
PlaywrightWebMCP
better · cheaperworse · pricier
← pricier than the fieldefficiency (σ, per task)cheaper than the field →

4 runs not plotted (missing a grade or token count — infra failures included).

Every metric, per harness × model

Benchmark pass rate

  1. claude-code / claude-haiku-4-5-20251001100% → 100%even
  2. codex / gpt-5.4-mini98% → 98%even

Task completion

  1. claude-code / claude-haiku-4-5-20251001100% → 100%even
  2. codex / gpt-5.4-mini100% → 100%even

Cost per run

  1. codex / gpt-5.4-mini$0.03 → $0.0230% better
  2. claude-code / claude-haiku-4-5-20251001$0.05 → $0.0343% better

Tokens per run

  1. codex / gpt-5.4-mini96.5k → 71k26% better
  2. claude-code / claude-haiku-4-5-20251001297.2k → 177.1k40% better

Duration per run

  1. claude-code / claude-haiku-4-5-2025100146s → 32.6s29% better
  2. codex / gpt-5.4-mini41.6s → 32.7s21% better

Playwright

webdriver

WebMCP

webmcp

The publisher has not described this setup — what each group's tools are, who built them, and how the tasks were chosen.

Distributions

Every completed run is one dot — the spread the averages hide. Click a dot to replay that run's journey.

Cost per run

PlaywrightWebMCP
code/claude-haiku-4-5-20251001
codex/gpt-5.4-mini
0$0.63

Token composition — average per run

InputOutputCache readCache write
code/claude-haiku-4-5-20251001 · base
504.8k
code/claude-haiku-4-5-20251001 · treat
182.5k
codex/gpt-5.4-mini · base
243k
codex/gpt-5.4-mini · treat
77k

Run outcomes

How each run ended — whether the agent finished, not whether its answer passed. Whether answers passed is in the results above.

CompletedFailureTimeoutError / other
code/claude-haiku-4-5-20251001 · base
50/50 completed
code/claude-haiku-4-5-20251001 · treat
50/50 completed
codex/gpt-5.4-mini · base
50/50 completed
codex/gpt-5.4-mini · treat
50/50 completed

Statistics

MetricArmnMeanMedian (pooled)MinMaxStd dev
Costbase100/100$0.07$0.04$0.01$0.63$0.10
treat100/100$0.03$0.02$0.01$0.07$0.01
Durationbase100/10059.6s42.6s24.4s4m 18s43.5s
treat100/10036.7s32.7s26.8s2m 00s12s
Total tokensbase100/100373.9k166.9k34.1k3.1M549.6k
treat100/100129.7k107.5k58k509.7k74.1k
Output tokensbase100/1001.7k92918815.6k2.2k
treat100/1007616453502.9k438
Cache-read tokensbase100/100350.6k148k21.8k3M527.4k
treat100/100121.3k101k32.9k500k75.1k
Cache-write tokensbase100/1008.9k2.2k081.6k14.9k
treat100/1002.9k1.7k036.7k4.4k
Turnsbase100/10015.25.529819.8
treat100/1008.373295.4

n = runs carrying the fact / completed runs in the arm; every statistic runs over present facts only — a missing fact is never counted as 0. Std dev is the sample form (n−1), withheld below n = 2. Every figure here POOLS all runs. The typical-run figures in the abstract, the head-to-head decision and the per-setup panels take each task's median first and then the median across tasks, so a task that ran more often never outweighs one that ran once; the two medians can differ. A promise study's headline (“cut average tokens”) compares per-arm means.

Task text withheld

Task text withheld — this benchmark is guarded and its tasks are not republished here.

Tasks

Runs

Every run is inspectable — open one to replay the agent's journey step by step, with the analyst's read underneath.

Every run, filterable… of 200Show runs
Task
Arm
Harness
Model
… of 200 runs match
Sort
HarnessModelTaskArmCompletedPassTokensCostDuration

Loading runs…

Methodology

What each metric means
Completed
The run finished and the harness returned a response. It does NOT mean the answer was correct — a completed run can score zero on quality.
Denominator: Terminal runs, excluding those killed by our own infrastructure.
Quality index
A composite used only in the quality x efficiency plot, built from whichever grader scored this study's runs — the judge's rubric, the pre-registered checks or the benchmark's own verdict — plus whether the run's own outcome was success. The map's caption names the grader and its weights.
Denominator: Graded runs — runs the study's grader scored.
Infra-excluded
A run killed by our own infrastructure. It is a missing measurement, never a loss for the arm, and is excluded from every rate denominator.
Denominator: Reported as a count beside every affected panel.
A/B arms
Every task runs twice per harness × model cell — once as Playwright, once as WebMCP — on the same prompt, same model, cold start for both arms. Both arms are real configurations.
One run per task
Every task ran once per cell and arm, in one phrasing. Run-to-run variation is therefore not measured: a single task's difference can be one lucky or unlucky attempt. Every row of the reference table treats the TASKS as the unit: its intervals resample the tasks and its p-values come from a paired test over them, so they describe how much the result depends on which tasks were drawn — not how it would change if the same tasks were run again. With few tasks no difference can reach significance (five tasks cannot go below p = 1/16).
Sample size
2 cells × 50 tasks × 2 arms = 200 runs.