Partial · Study2views

benchme — in-pod Playwright vs Kitesurf (50 tasks)

Kitesurf (Playwright) vs Playwright (in-pod) · 50 tasks · 2 harnesses · 2 models

Playwright (in-pod) is the better choice over Kitesurf (Playwright) — cost +18%, tokens +27.6%, quality +1.7pts, with the goal gain concentrated in one cell rather than uniform.

Correctness did not separate the arms beyond the decision margin, and the baseline got there with 54% faster and 22% fewer tokens.

Abstract

Across 50 tasks over 2 harness/model cells, Kitesurf (Playwright) ran more expensively than Playwright (in-pod) — cost +18%, a median of $0.05 against $0.04 (means $0.05 against $0.05); tokens +27.6%, wall-clock +118.5%. The tail runs the other way: the treatment's worst case is worse than the baseline's. Quality improved: rubric score 88/100 against 87/100 (+1.7pts), and blind pairwise judging found 91% equivalent, 4% treatment better, 5% baseline better (of 100 judged pairs). Deterministic checks passed more often: 98% against 96% (+2pp (+2.1% relative)). Goal achievement moved 48% to 52%, +4pp (+8.3% relative). That gain is concentrated rather than general — codex/gpt-5.4-mini drives it (+4pp pooled; per-cell +2pp to +6pp). Harness and model are not crossed in this matrix, so neither can be credited alone.

The result

Playwright (in-pod)Kitesurf (Playwright)

Best arm overall: Playwright (in-pod) (72.3 vs 68.4 of 100)

Best setup: claude-code / claude-haiku-4-5-20251001 · Playwright (in-pod) (84.1 of 100)

  1. 1claude-code / claude-haiku-4-5-20251001Playwright (in-pod)84.1best
  2. 2claude-code / claude-haiku-4-5-20251001Kitesurf (Playwright)79.8
  3. 3codex / gpt-5.4-miniPlaywright (in-pod)65.0
  4. 4codex / gpt-5.4-miniKitesurf (Playwright)62.0

Overall score per arm, 0–100 points (not a pass rate): 30% goal + 30% quality + 15% evals + 25% efficiency (cost · tokens · duration, vs the board's best arm). Absent components renormalize. Best at the top — the board's best setup is tagged.

Head-to-head

Correctness did not separate the arms beyond the decision margin, and the baseline got there with 54% faster and 22% fewer tokens.

Same result both sides — decided on what it cost

What the blind judge alone said

Playwright (in-pod) wins 5Ties 91Treatment wins 4Win rate 4%

In blind position-debiased comparison, Kitesurf (Playwright) won 4 of 100 pairs against Playwright (in-pod); 91 were ties.

Quality × efficiency clusters — normalized per task

Every graded run, standardized WITHIN its task so difficulty cancels out: → right = fewer tokens than the field on the same task, ↑ up = higher quality index (70% rubric pass rate + 15% intent + 15% outcome) than the field. The field pools every harness and model, so a setup's left–right position largely reflects its own token habits. Dots are runs; each harness logo is a harness × model × arm centroid — hover it for the model and averages. Up-right wins.

Label
Task
Arm
Harness
Model
Verdict
Evals
196 of 196 runs match
Show
Playwright (in-pod)Kitesurf (Playwright)
better · cheaperworse · pricier
← pricier than the fieldefficiency (σ, per task)cheaper than the field →

4 runs not plotted (missing a grade or token count — infra failures included).

Every metric, per harness × model

Task completion

  1. claude-code / claude-haiku-4-5-20251001100% → 100%even
  2. codex / gpt-5.4-mini100% → 100%even

Goal achievement

  1. claude-code / claude-haiku-4-5-2025100188% → 90%2.0 pp better
  2. codex / gpt-5.4-mini8% → 14%6.0 pp better

Rubric quality /100

  1. claude-code / claude-haiku-4-5-2025100196 → 982.1 pts better
  2. codex / gpt-5.4-mini77 → 79even

Evals pass rate

  1. claude-code / claude-haiku-4-5-2025100196% → 98%2.0 pp better
  2. codex / gpt-5.4-mini96% → 98%2.0 pp better

Cost per run

  1. codex / gpt-5.4-mini$0.03 → $0.038% worse
  2. claude-code / claude-haiku-4-5-20251001$0.05 → $0.0626% worse

Tokens per run

  1. codex / gpt-5.4-mini88.2k → 102k16% worse
  2. claude-code / claude-haiku-4-5-20251001274.4k → 376k37% worse

Duration per run

  1. codex / gpt-5.4-mini30.9s → 55.8s80% worse
  2. claude-code / claude-haiku-4-5-2025100136.5s → 1m 27s139% worse

Rubric criteria passed

  1. claude-code / claude-haiku-4-5-2025100196% → 98%2.1 pp better
  2. codex / gpt-5.4-mini77% → 78%even

Playwright (in-pod)

webdriver

Kitesurf (Playwright)

kitesurf

Deterministic checks

Pre-registered pass/fail expectations — MCP calls, commands, files, responses — checked deterministically against each run's recorded evidence, independent of any LLM judge.

Pass rate — Playwright (in-pod) vs Kitesurf (Playwright)96% vs 98% · +2ppShow per task (50)
Playwright (in-pod)96%
Kitesurf (Playwright)98%
Δ pass rate+2pp
T1vaultdocsPlaywright (in-pod)2/2Kitesurf (Playwright)2/2
T2warehousePlaywright (in-pod)2/2Kitesurf (Playwright)2/2
T3warehousePlaywright (in-pod)2/2Kitesurf (Playwright)2/2
T4warehousePlaywright (in-pod)2/2Kitesurf (Playwright)2/2
T5warehousePlaywright (in-pod)2/2Kitesurf (Playwright)2/2
T6warehousePlaywright (in-pod)2/2Kitesurf (Playwright)2/2
T7warehousePlaywright (in-pod)2/2Kitesurf (Playwright)2/2
T8warehousePlaywright (in-pod)1/2Kitesurf (Playwright)2/2
T9warehousePlaywright (in-pod)2/2Kitesurf (Playwright)2/2
T10warehousePlaywright (in-pod)2/2Kitesurf (Playwright)2/2
T11warehousePlaywright (in-pod)2/2Kitesurf (Playwright)2/2
T12warehousePlaywright (in-pod)2/2Kitesurf (Playwright)2/2
T13warehousePlaywright (in-pod)2/2Kitesurf (Playwright)2/2
T14warehousePlaywright (in-pod)2/2Kitesurf (Playwright)2/2
T15warehousePlaywright (in-pod)2/2Kitesurf (Playwright)2/2
T16warehousePlaywright (in-pod)2/2Kitesurf (Playwright)2/2
T17warehousePlaywright (in-pod)2/2Kitesurf (Playwright)2/2
T18warehousePlaywright (in-pod)2/2Kitesurf (Playwright)2/2
T19warehousePlaywright (in-pod)2/2Kitesurf (Playwright)2/2
T20warehousePlaywright (in-pod)2/2Kitesurf (Playwright)2/2
T21warehousePlaywright (in-pod)2/2Kitesurf (Playwright)2/2
T22warehousePlaywright (in-pod)2/2Kitesurf (Playwright)2/2
T23warehousePlaywright (in-pod)2/2Kitesurf (Playwright)2/2
T24warehousePlaywright (in-pod)2/2Kitesurf (Playwright)2/2
T25warehousePlaywright (in-pod)2/2Kitesurf (Playwright)2/2
T26warehousePlaywright (in-pod)2/2Kitesurf (Playwright)2/2
T27warehousePlaywright (in-pod)0/2Kitesurf (Playwright)0/2
T28helpdeskPlaywright (in-pod)2/2Kitesurf (Playwright)2/2
T29helpdeskPlaywright (in-pod)2/2Kitesurf (Playwright)2/2
T30helpdeskPlaywright (in-pod)2/2Kitesurf (Playwright)2/2
T31helpdeskPlaywright (in-pod)2/2Kitesurf (Playwright)2/2
T32helpdeskPlaywright (in-pod)2/2Kitesurf (Playwright)2/2
T33helpdeskPlaywright (in-pod)2/2Kitesurf (Playwright)2/2
T34helpdeskPlaywright (in-pod)2/2Kitesurf (Playwright)2/2
T35helpdeskPlaywright (in-pod)2/2Kitesurf (Playwright)2/2
T36helpdeskPlaywright (in-pod)2/2Kitesurf (Playwright)2/2
T37helpdeskPlaywright (in-pod)2/2Kitesurf (Playwright)2/2
T38helpdeskPlaywright (in-pod)2/2Kitesurf (Playwright)2/2
T39helpdeskPlaywright (in-pod)2/2Kitesurf (Playwright)2/2
T40helpdeskPlaywright (in-pod)1/2Kitesurf (Playwright)2/2
T41vaultdocsPlaywright (in-pod)2/2Kitesurf (Playwright)2/2
T42vaultdocsPlaywright (in-pod)2/2Kitesurf (Playwright)2/2
T43vaultdocsPlaywright (in-pod)2/2Kitesurf (Playwright)2/2
T44vaultdocsPlaywright (in-pod)2/2Kitesurf (Playwright)2/2
T45vaultdocsPlaywright (in-pod)2/2Kitesurf (Playwright)2/2
T46vaultdocsPlaywright (in-pod)2/2Kitesurf (Playwright)2/2
T47vaultdocsPlaywright (in-pod)2/2Kitesurf (Playwright)2/2
T48vaultdocsPlaywright (in-pod)2/2Kitesurf (Playwright)2/2
T49vaultdocsPlaywright (in-pod)2/2Kitesurf (Playwright)2/2
T50vaultdocsPlaywright (in-pod)2/2Kitesurf (Playwright)2/2

Distributions

Every completed run is one dot — the spread the averages hide. Click a dot to replay that run's journey.

Rubric quality per run

Playwright (in-pod)Kitesurf (Playwright)
code/claude-haiku-4-5-20251001
codex/gpt-5.4-mini
0100

Token composition — average per run

InputOutputCache readCache write
code/claude-haiku-4-5-20251001 · base
333.1k
code/claude-haiku-4-5-20251001 · treat
420.3k
codex/gpt-5.4-mini · base
96.9k
codex/gpt-5.4-mini · treat
120.5k

Where agents fail

Each cell is the pass fraction of one pre-registered criterion — darker red, more failures. Click a fraction to read the judge's notes on the runs behind it.

Criterionclaude-codecodexAll
T2The response text itself (not the run trace) identifies the product record the price comes from, e.g. by the product name and an identifier the warehouse app shows, so a reader can tell which record the figure relies on.0/2
T9The response text itself (not the run trace) identifies the product record the price comes from, e.g. by the product name and an identifier the warehouse app shows, so a reader can tell which record the figure relies on.0/2
T26The response states the basis for its pick: the unit price it found for the chosen product and the prices it compared it against, rather than a bare SKU with no stated basis.0/2
T27The "city" value in the final JSON object is the string "Tessalind" (case-insensitive, surrounding whitespace ignored).0/2
T49The response shows its basis: it names the document it read the figure from and identifies it as the final FY2025 results summary, not the superseded draft the task warns about.0/2
T1The response text itself (not the run trace) names the document the id belongs to, e.g. its title as the vault shows it, so a reader can see which record the answer relies on.1/2
T2The response text presents the figure as the unit price of the Quenquen Renmar kit, the product the task names, and not as another product's price or a total for several units.1/2
T3The response answers the request as asked: the value it reports is a customer code (not only a customer name or another identifier), and the response text presents that customer as gold-tier and based in Quillhaven, addressing each condition the task set.1/2
T3The response text itself (not the run trace) identifies the customer record the code came from, e.g. the customer's name as the warehouse app shows it, so a reader can see which record the answer relies on.1/2
T4The response text presents the figure as the current in-stock quantity of the Falnexmar Kael assembly (SKU ENC-1107) across all depots, as asked, and not as the count at an individual depot or a figure for another item.1/2
T4The response text itself (not the run trace) says which item record in the warehouse app the figure comes from, so a reader can trace where the number came from.1/2
T5The response text itself (not the run trace) says which item record in the warehouse app the figure comes from, so a reader can trace where the number came from.1/2
T6The response text presents the figure as the current in-stock quantity of the Zyom Renom kit (SKU FAS-1000) across all depots, as asked, and not as the count at an individual depot or a figure for another item.1/2
T6The response text itself (not the run trace) says which item record in the warehouse app the figure comes from, so a reader can trace where the number came from.1/2
T7The response text itself (not the run trace) identifies the item and the depot entry the figure was read from, so a reader can trace where the number came from.1/2
T8The response text itself (not the run trace) identifies the item and the depot entry the figure was read from, so a reader can trace where the number came from.1/2
T9The response text presents the figure as the unit price of the Lolo Vivi pack, the product the task names, and not as another product's price or a total for several units.1/2
T10The response text presents the figure as the unit price of the Vizynex Renquen pack, the product the task names, and not as another product's price or a total for several units.1/2
T10The response text itself (not the run trace) identifies the product record the price comes from, e.g. by the product name and an identifier the warehouse app shows, so a reader can tell which record the figure relies on.1/2
T11The response shows its basis: it names the customer the code belongs to and states that this customer is platinum-tier and based in Karrow, addressing the conditions the task set.1/2
T12The response names the customer the task asked about (Quenfalsil Industrial) together with the tier, so the value is clearly tied to that customer and not another.1/2
T13The response names the customer the task asked about (Kasilvi Systems) together with the tier, so the value is clearly tied to that customer and not another.1/2
T14The response makes clear the count is of the open orders of Vitru Industrial (C-127) and shows its basis: it identifies the orders it counted (e.g. by order number) or says how it decided which of that customer's orders were open.1/2
T15The response makes clear the count is of the open orders of Thomarom Systems (C-115) and shows its basis: it identifies the orders it counted (e.g. by order number) or says how it decided which of that customer's orders were open.1/2
T16The response names the order it looked up (SO-20003) alongside the status, so the value is clearly tied to that order and not another.1/2
T17The response names the order it looked up (SO-20040) alongside the status, so the value is clearly tied to that order and not another.1/2
T18The response names the order it looked up (SO-20167) alongside the status, so the value is clearly tied to that order and not another.1/2
T19The response shows the basis for the total: it names the order it read (SO-20296) and either lists the line figures it summed or says which of that order's own figures it relied on.1/2
T20The response shows the basis for the total: it names the order it read (SO-20102) and either lists the line figures it summed or says which of that order's own figures it relied on.1/2
T21The response identifies its value as the total of order SO-20156 across all of that order's lines (it names that order), not as a single line's amount or another order's figure.1/2
T22The response identifies its value as the quantity of SKU ENC-1075 on order SO-20060 (it names that SKU or order), not as a different figure such as the order's total units or stock on hand.1/2
T23The response identifies its value as the quantity of SKU CAB-1034 on order SO-20118 (it names that SKU or order), not as a different figure such as the order's total units or stock on hand.1/2
T24The response identifies the code as the destination of transfer TR-532 (it names that transfer), not as that transfer's origin or another transfer's destination.1/2
T25The response frames its answer as the depot with the most units of SKU FLU-1055 specifically (it names that SKU or product), not as the depot with the most stock overall.1/2
T25The response states the basis for its choice: the stock figures for SKU FLU-1055 it compared across depots, rather than a bare depot code with no stated basis.1/2
T26The response frames its pick as the enclosures-category product with the highest unit price (it states unit price as the ranking criterion), not as a pick by stock level, order volume or popularity.1/2
T27The response identifies the city as the location of depot QHV (it names that depot by code or name), not as the city of a different depot.1/2
T31The response presents the value as the priority of ticket HD-5128 itself (it names that ticket), not the ticket's status or the priority of a different ticket.1/2
T32The response presents the value as the priority of ticket HD-5136 itself (it names that ticket), not the ticket's status or the priority of a different ticket.1/2
T35The response presents the agent as the assignee of ticket HD-5025 (it names that ticket), not the ticket's requester (the customer) or the agent on a different ticket.1/2
T38The response presents the value as the team of support agent Lior Kowalczyk (it names that agent), not a different agent's team.1/2
T39The response presents the value as the team of support agent Quill Ferrante (it names that agent), not a different agent's team.1/2
T40The response presents the number as the count of helpdesk tickets reported by customer Omthonex Supply (C-131) (it names that customer), not tickets handled by an agent or another customer's tickets; if it counted only some of that customer's tickets, it says so.1/2
T40The response states the basis for its count: it identifies the tickets it counted (e.g. by ticket number) or the record or listing it took the figure from, rather than giving a bare number.1/2
T41The response ties the id to the document the task asked about: alongside the id it names that document (e.g. its title as shown in the vault), and the document it names is operations memo 2: rentruzy retrospective, not a different memo.1/2
T42The response labels the value it reports as the standard supplier lead time for the Zyom Renom kit (SKU FAS-1000) — it names that product or SKU — not as a different kind of figure or another product's figure.1/2
T42The response shows its basis: it names the vault document it read the figure from (e.g. its title as shown in the vault), so a reader can see that the answer relies on the datasheet for SKU FAS-1000.1/2
T43The response labels the value it reports as the standard supplier lead time for the Lofalel Zynex assembly (SKU ADH-1054) — it names that product or SKU — not as a different kind of figure or another product's figure.1/2
T43The response shows its basis: it names the vault document it read the figure from (e.g. its title as shown in the vault), so a reader can see that the answer relies on the datasheet for SKU ADH-1054.1/2
T44The response labels the value it reports as the minimum order quantity for the Baelfal Marren pack (SKU CAB-1018) — it names that product or SKU — not as a different kind of figure or another product's figure.1/2
T44The response shows its basis: it names the vault document it read the figure from (e.g. its title as shown in the vault), so a reader can see that the answer relies on the datasheet for SKU CAB-1018.1/2
T45The response labels the value it reports as the discount off list price for platinum customers under the customer tier discounts policy (it names the platinum tier), not as another tier's discount or a different term of that policy.1/2
T45The response shows its basis: it names the vault document it read the figure from (e.g. its title as shown in the vault), so a reader can see that the answer relies on the customer tier discounts policy.1/2
T46The response labels the value it reports as the number of days after delivery within which customers may return unused goods, not as a time limit for something else or a figure in other units.1/2
T46The response shows its basis: it names the vault document it read the figure from (e.g. its title as shown in the vault), so a reader can see that the answer relies on the returns and warranty policy.1/2
T47The response labels the hour it reports as the time until which Thoomsil depot 3 (QHV) dispatches goods — it names that depot by name or code — not as some other hour of that depot or another depot's hour.1/2
T47The response shows its basis: it names the vault document it read the figure from (e.g. its title as shown in the vault), so a reader can see that the answer relies on the depot operating hours policy.1/2
T48The response labels the value it reports as the number of months until Trumar Systems' next review — it names that account — not as a calendar date, a duration in other units, or another account's review timing.1/2
T48The response shows its basis: it names the vault document it read the figure from (e.g. its title as shown in the vault), so a reader can see that the answer relies on the account review notes for Trumar Systems.1/2
T49The response labels the value it reports as the company's number of employees for FY2025, not as a different kind of figure or a figure for another period.1/2
T50The response labels the value it reports as the company's FY2025 revenue in US dollars, not as a different kind of figure, a figure in another currency, or a figure for another period.1/2
T50The response shows its basis: it names the document it read the figure from and identifies it as the final FY2025 results summary, not the superseded draft the task warns about.1/2
62 of 283 graded criteria failed at least once.

Run outcomes

How each run ended — whether the agent finished, not whether its answer passed. Whether answers passed is in the results above.

CompletedFailureTimeoutError / other
code/claude-haiku-4-5-20251001 · base
50/50 completed
code/claude-haiku-4-5-20251001 · treat
50/50 completed
codex/gpt-5.4-mini · base
50/50 completed
codex/gpt-5.4-mini · treat
50/50 completed

Efficiency frontier

Label
Task
Arm
Harness
Model
Verdict
Evals
196 of 196 runs match
Loading
Playwright (in-pod)Kitesurf (Playwright)
claude-codecodex

Statistics

MetricArmnMeanMedian (pooled)MinMaxStd dev
Rubric qualitybase100/10087831710015
treat100/100881006010013
Costbase98/100$0.05$0.03$0.02$0.30$0.04
treat98/100$0.05$0.05$0.02$0.37$0.04
Durationbase100/10035.9s32.8s24.4s1m 47s11.3s
treat100/1001m 19s1m 13s44.8s4m 09s31.6s
Total tokensbase98/100212.6k157.6k62.2k1.8M230.6k
treat98/100270.4k240k77.5k2.1M250k
Output tokensbase98/1001.1k9324525.5k772
treat98/1001.3k1.1k5606.8k835
Cache-read tokensbase98/100195.3k149.4k47.6k1.7M222.1k
treat98/100253.3k229.6k59.8k2M240k
Cache-write tokensbase98/1009.5k0081.7k14k
treat98/1009k3.8k0105.2k13.9k
Turnsbase100/100129.537411.4
treat100/10015.21038313.5

n = runs carrying the fact / completed runs in the arm; every statistic runs over present facts only — a missing fact is never counted as 0. Std dev is the sample form (n−1), withheld below n = 2. Every figure here POOLS all runs. The typical-run figures in the abstract, the head-to-head decision and the per-setup panels take each task's median first and then the median across tasks, so a task that ran more often never outweighs one that ran once; the two medians can differ. A promise study's headline (“cut average tokens”) compares per-arm means.

Tasks

Runs

Every run is inspectable — open one to replay the agent's journey step by step, with the analyst's read underneath.

Every run, filterable… of 200Show runs
Task
Arm
Harness
Model
… of 200 runs match
Sort
HarnessModelTaskArmCompletedQualityEvalsPairTokensCostDuration

Loading runs…

Methodology

What each metric means
Completed
The run finished and the harness returned a response. It does NOT mean the answer was correct — a completed run can score zero on quality.
Denominator: Terminal runs, excluding those killed by our own infrastructure.
Goal achievement
The run fully achieved the task's goal: the judge panel passed every pre-registered rubric criterion (a rubric score of 100).
Denominator: Graded completed runs.
Rubric quality
The share of pre-registered acceptance criteria the judge panel passed, as a 0-100 score. Partial credit is possible.
Denominator: The criteria count frozen before any run — never the judge's returned count.
Quality index
A composite used only in the quality x efficiency plot, built from whichever grader scored this study's runs — the judge's rubric, the pre-registered checks or the benchmark's own verdict — plus whether the run's own outcome was success. The map's caption names the grader and its weights.
Denominator: Graded runs — runs the study's grader scored.
Pairwise verdict
A blind, position-debiased comparison of the two arms' final answers. It sees answer text only — cost and latency are measured separately.
Denominator: Task-cell pairs where both arms produced a response.
Infra-excluded
A run killed by our own infrastructure. It is a missing measurement, never a loss for the arm, and is excluded from every rate denominator.
Denominator: Reported as a count beside every affected panel.
Eval pass rate
The share of pre-registered deterministic expectations (MCP calls, commands, files, responses) that passed for an arm — checked mechanically against recorded evidence, independent of the LLM judge. Changes between arms are read in PERCENTAGE POINTS, the same convention as every other rate on this page.
Denominator: Runs with a MEASURABLE eval score (unmeasurable and not-yet-graded runs are excluded, never counted as failures).
Abstained
A check whose evidence channel this harness cannot produce (a deny-by-default capability roster) — NEITHER a pass NOR a fail, and never folded into the failure count.
Denominator: Reported as its own count beside every eval readout it affects.
A/B arms
Every task runs twice per harness × model cell — once as Playwright (in-pod), once as Kitesurf (Playwright) — on the same prompt, same model, cold start for both arms. Both arms are real configurations; the judges never see which is which.
Pre-registered rubrics
Each task's acceptance criteria (5–6 per task) are written at task generation, before any run exists, so grading can never be shaped by the results.
Judge panel
A panel of independent judges (claude-opus-5, gpt-5.5, claude-fable-5), each at provider-default sampling (claude-opus-5, gpt-5.5 and claude-fable-5 do not accept a temperature setting), scores every successful response against its rubric; a criterion passes only when a strict majority of the panel passes it. Each judge sees only the task, the rubric, and the response — never which arm produced it, never token counts.
Blind pairwise
Each task's two responses are also compared blind as “Response A” and “Response B” by every panel judge, each judging twice with the order swapped (disagreement = tie); the pair's verdict is the panel's strict majority — no majority counts as a tie.
How the head-to-head is decided
The conclusion is not the blind judge's alone — that judge sees only the two response texts, never the rubric score or what each run cost. Evidence is ranked in order: whether the task's goal was achieved, then the rubric score, then the measured cost of getting there (tokens, turns, duration, spend), and only then the blind judge, for pairs nothing else separates. Cost never outranks correctness, and when neither arm achieved the goal the efficiency gap between them decides nothing.
Deterministic eval checks
Some tasks also carry pre-registered, deterministic expectations — an MCP tool called with particular arguments, a command run, a file touched, a response containing or matching a pattern — checked mechanically against the run's recorded evidence, with no model in the loop. A check whose evidence this harness cannot produce abstains rather than failing. When this lens is enabled, it slots into the evidence order above right after whether the task's goal was achieved and before the rubric score.
One run per task
Every task ran once per cell and arm, in one phrasing. Run-to-run variation is therefore not measured: a single task's difference can be one lucky or unlucky attempt. Every row of the reference table treats the TASKS as the unit: its intervals resample the tasks and its p-values come from a paired test over them, so they describe how much the result depends on which tasks were drawn — not how it would change if the same tasks were run again. With few tasks no difference can reach significance (five tasks cannot go below p = 1/16).
Sample size
2 cells × 50 tasks × 2 arms = 200 runs.