Not delivered · Study1views
Web Research study
With Web Research vs Baseline · 5 tasks · 2 harnesses · 2 models
Web Research does not justify its efficiency cost.
Both arms reached the same result, and the baseline got there with 36% fewer turns and 30% lower cost.
Abstract
Web Research increased costs and latency across 5 tasks over 2 harness/model cells. Cost changed +43.8%, with a median of $0.06 (mean $0.05) versus $0.04 (mean $0.05) for baseline. Token consumption changed +29.2%, reaching a median of 133k (mean 135.7k) against 103k (mean 102k). Duration increased +25.2%, with a median of 1m 00s (mean 1m 19s) compared to 48s (mean 56.6s). The treatment's worst case is worse than the baseline's: duration extended to 3m 27s against 2m 14s, and tokens reached 236.5k against 188.6k. Because harness and model are not crossed, effects cannot be attributed to either alone.
The result
BaselineWith Web Research
Best arm overall: Baseline (69.2 vs 64.0 of 100)
Same outcome on both arms — the whole gap is efficiency (cost · tokens · duration).
Best setup: claude-code / claude-haiku-4-5-20251001 · Baseline (74.2 of 100)
- 1claude-code / claude-haiku-4-5-20251001Baseline74.2best
- 2claude-code / claude-haiku-4-5-20251001With Web Research68.1
- 3codex / gpt-5.4-miniBaseline65.0
- 4codex / gpt-5.4-miniWith Web Research61.0
Overall score per arm, 0–100 points (not a pass rate): 30% goal + 30% quality + 25% efficiency (cost · tokens · duration, vs the board's best arm). Absent components renormalize. Best at the top — the board's best setup is tagged.
Head-to-head
Both arms reached the same result, and the baseline got there with 36% fewer turns and 30% lower cost.
Same result both sides — decided on what it cost
What the blind judge alone said
Baseline wins 1Ties 4Web Research wins 5Win rate 50%
In blind position-debiased comparison, Web Research won 5 of 10 pairs; 4 were ties.
Quality × efficiency clusters — normalized per task
Every graded run, standardized WITHIN its task so difficulty cancels out: → right = fewer tokens than the field on the same task, ↑ up = higher quality index (70% rubric pass rate + 15% intent + 15% outcome) than the field. The field pools every harness and model, so a setup's left–right position largely reflects its own token habits. Dots are runs; each harness logo is a harness × model × arm centroid — hover it for the model and averages. Up-right wins.
Every metric, per harness × model
Task completion
- claude-code / claude-haiku-4-5-20251001100% → 100%even
- codex / gpt-5.4-mini100% → 100%even
Goal achievement
- claude-code / claude-haiku-4-5-2025100160% → 60%even
- codex / gpt-5.4-mini40% → 40%even
Rubric quality /100
- claude-code / claude-haiku-4-5-2025100184 → 884.0 pts better
- codex / gpt-5.4-mini84 → 804.0 pts worse
Cost per run
- claude-code / claude-haiku-4-5-20251001$0.03 → $0.0451% worse
- codex / gpt-5.4-mini$0.07 → $0.0812% worse
Tokens per run
- codex / gpt-5.4-mini59.9k → 70.7k18% worse
- claude-code / claude-haiku-4-5-20251001145.8k → 225.8k55% worse
Duration per run
- codex / gpt-5.4-mini1m 03s → 1m 09s10% worse
- claude-code / claude-haiku-4-5-2025100149.4s → 1m 09s40% worse
Rubric criteria passed
- claude-code / claude-haiku-4-5-2025100184% → 88%4.0 pp better
- codex / gpt-5.4-mini84% → 80%4.0 pp worse
Baseline
No add-ons — the agent's built-in tools only.
With Web Research
No add-ons — the agent's built-in tools only.
The publisher has not described this setup — what each group's tools are, who built them, and how the tasks were chosen.
Distributions
Every completed run is one dot — the spread the averages hide. Click a dot to replay that run's journey.
Rubric quality per run
Token composition — average per run
Where agents fail
Each cell is the pass fraction of one pre-registered criterion — darker red, more failures. Click a fraction to read the judge's notes on the runs behind it.
| Criterion | claude-code | codex | All |
|---|---|---|---|
| T5Response identifies the 2 comparison libraries by name and cites sources that benchmark them against FastCache. | 0/2 | ||
| T5Response cites at least 2 independent sources (e.g., published benchmark reports, GitHub issue discussions, academic papers) that compare the libraries. | 0/2 | ||
| T1Response distinguishes between direct purchase price and cloud rental pricing, or explains why only one applies. | 1/2 | ||
| T4Response cites the official GDPR text (e.g., EUR-Lex or the official EU GDPR website) as a source. | 1/2 | ||
| T4Each of the 3 obligations is accompanied by a brief explanation of what it means in practice for a SaaS platform. | 1/2 | ||
| T4Response distinguishes between what Article 25 explicitly states and any inferences about how it applies (e.g., 'Article 25 requires X; this implies Y for SaaS platforms'). | 1/2 |
Run outcomes
How each run ended — whether the agent finished, not whether its answer passed. Whether answers passed is in the results above.
Efficiency frontier
Statistics
| Metric | Arm | n | Mean | Median (pooled) | Min | Max | Std dev |
|---|---|---|---|---|---|---|---|
| Rubric quality | base | 10/10 | 84 | 90 | 60 | 100 | 18 |
| treat | 10/10 | 84 | 90 | 60 | 100 | 18 | |
| Cost | base | 10/10 | $0.05 | $0.03 | $0.02 | $0.12 | $0.03 |
| treat | 10/10 | $0.06 | $0.05 | $0.03 | $0.11 | $0.03 | |
| Duration | base | 10/10 | 56.6s | 50.4s | 26.7s | 2m 14s | 31.3s |
| treat | 10/10 | 1m 19s | 1m 09s | 32.9s | 3m 27s | 51.4s | |
| Total tokens | base | 10/10 | 102k | 90.9k | 25k | 188.6k | 59.4k |
| treat | 10/10 | 135.7k | 101.2k | 35.4k | 236.5k | 83.8k | |
| Output tokens | base | 10/10 | 3.3k | 2k | 900 | 9.4k | 2.9k |
| treat | 10/10 | 4.6k | 3.2k | 1.8k | 11.2k | 3.2k | |
| Cache-read tokens | base | 10/10 | 66.8k | 37.7k | 0 | 178.4k | 75.1k |
| treat | 10/10 | 96.9k | 50.6k | 0 | 220.7k | 107.8k | |
| Cache-write tokens | base | 10/10 | 6k | 2.3k | 0 | 37.2k | 11.4k |
| treat | 10/10 | 4.8k | 4.2k | 0 | 12.4k | 5.2k | |
| Turns | base | 10/10 | 6.2 | 4 | 2 | 15 | 4.9 |
| treat | 10/10 | 10.2 | 7.5 | 2 | 20 | 8.8 |
n = runs carrying the fact / completed runs in the arm; every statistic runs over present facts only — a missing fact is never counted as 0. Std dev is the sample form (n−1), withheld below n = 2. Every figure here POOLS all runs. The typical-run figures in the abstract, the head-to-head decision and the per-setup panels take each task's median first and then the median across tasks, so a task that ran more often never outweighs one that ran once; the two medians can differ. A promise study's headline (“cut average tokens”) compares per-arm means.
Tasks
What is the current price of NVIDIA's H100 GPU as of today, and which major cloud providers offer it? List at least 3 cloud providers with their current pricing or availability status.
Tests whether the skill's decomposition and primary-source lookup methods retrieve accurate, current pricing information across multiple vendors. Guessing product prices is a high-cost failure mode the skill directly addresses.
Pre-registered rubric
- Response lists at least 3 major cloud providers (e.g., AWS, Google Cloud, Azure, or equivalent) that offer H100 access.
- For each provider, response includes either a specific current price/cost or a clear statement that pricing is not publicly listed.
- Response cites at least 2 distinct sources (URLs or vendor pages) for the pricing or availability claims.
- Response distinguishes between direct purchase price and cloud rental pricing, or explains why only one applies.
- Response includes a publication date or freshness indicator (e.g., 'as of [date]' or 'last checked [date]') for at least one claim.
The Rust programming language just released a new stable version. Find the release notes and identify 3 significant new features or changes introduced in this latest stable release. For each feature, cite the source.
Tests decomposition into sub-questions (What is the latest version? What are its release notes? What are the significant features?) and primary-source lookup (official Rust release channels). The skill's emphasis on vendor docs over aggregators is critical here.
Pre-registered rubric
- Response identifies a specific stable Rust version number (e.g., 1.X.Y).
- Response lists exactly 3 distinct new features or breaking changes introduced in that version.
- Each of the 3 features is accompanied by a citation to the official Rust release notes URL or rust-lang.org domain.
- Response includes the release date of the identified version.
- For at least one feature, the response includes a brief explanation of its impact or use case (not just a name).
A software engineer claims that 'Python's garbage collector causes a 20% performance overhead on CPU-bound workloads.' Research this claim: is it accurate, exaggerated, or context-dependent? Find at least 2 independent sources that address Python GC performance impact, and explain whether they support, refute, or qualify the claim.
Tests the skill's cross-checking and adversarial verification methods. A simplistic agent might return the first search result; the skill requires reading primary sources, comparing claims, and calibrated uncertainty. This is exactly the scenario where the skill's 'never fabricate, distinguish inference from fact' rules prevent misleading answers.
Pre-registered rubric
- Response cites at least 2 distinct, independent sources (e.g., academic papers, official Python docs, blog posts by recognized experts).
- For each source, response clearly states what it claims about Python GC performance (e.g., a specific percentage, a conditional statement, or a refutation).
- Response explicitly indicates whether the sources agree or disagree with the original claim.
- Response distinguishes between the claim as stated ('20% overhead') and any nuance from sources (e.g., 'overhead depends on workload type' or 'only applies to certain Python implementations').
- Response indicates whether the engineer's claim is accurate, exaggerated, context-dependent, or unsupported, with a brief justification based on the sources.
A regulatory requirement states that SaaS platforms handling EU customer data must comply with GDPR Article 25. Research what Article 25 specifically requires and summarize the 3 most critical obligations for a SaaS platform. Cite primary sources.
Tests the skill's ability to locate and correctly interpret regulatory primary sources (official GDPR text, not blog summaries). Misreading regulations is a high-stakes failure mode. The skill's emphasis on reading actual pages and cross-checking against multiple sources prevents mis-citation.
Pre-registered rubric
- Response correctly identifies the title or subject of GDPR Article 25 (e.g., 'Data Protection by Design and by Default').
- Response lists exactly 3 distinct obligations from Article 25, expressed in concrete terms (not generic compliance language).
- Response cites the official GDPR text (e.g., EUR-Lex or the official EU GDPR website) as a source.
- Each of the 3 obligations is accompanied by a brief explanation of what it means in practice for a SaaS platform.
- Response distinguishes between what Article 25 explicitly states and any inferences about how it applies (e.g., 'Article 25 requires X; this implies Y for SaaS platforms').
An open-source library 'FastCache' claims to be 'the fastest caching library for Python.' Research independent benchmarks or comparisons between FastCache and at least 2 other major Python caching libraries (e.g., Cachetools, Redis-py, Diskcache). Report what the benchmarks show about relative performance.
Tests decomposition (Which caching libraries are major competitors? Where are independent benchmarks published?) and primary-source reading (finding actual benchmark reports, not marketing claims). The skill's rule about reading vendor docs vs. aggregators applies: distinguishing independent benchmarks from vendor hype.
Pre-registered rubric
- Response identifies the 2 comparison libraries by name and cites sources that benchmark them against FastCache.
- Response reports specific metrics from the benchmarks (e.g., throughput, latency, memory usage) rather than vague performance claims.
- Response cites at least 2 independent sources (e.g., published benchmark reports, GitHub issue discussions, academic papers) that compare the libraries.
- Response notes the benchmark conditions (e.g., data size, workload type) that apply to the reported performance metrics.
- Response indicates whether FastCache's claim to be 'fastest' is supported, refuted, or qualified by the benchmarks (e.g., 'fastest under condition X but not Y').
Runs
Every run is inspectable — open one to replay the agent's journey step by step, with the analyst's read underneath.
Every run, filterable… of 20Show runsHide runs
| Harness | Model | Task | Arm | Completed | Quality | Pair | Tokens | Cost | Duration |
|---|
Loading runs…
Methodology
What each metric means
- Completed
- The run finished and the harness returned a response. It does NOT mean the answer was correct — a completed run can score zero on quality.
- Denominator: Terminal runs, excluding those killed by our own infrastructure.
- Goal achievement
- The run fully achieved the task's goal: the judge panel passed every pre-registered rubric criterion (a rubric score of 100).
- Denominator: Graded completed runs.
- Rubric quality
- The share of pre-registered acceptance criteria the judge panel passed, as a 0-100 score. Partial credit is possible.
- Denominator: The criteria count frozen before any run — never the judge's returned count.
- Quality index
- A composite used only in the quality x efficiency plot, built from whichever grader scored this study's runs — the judge's rubric, the pre-registered checks or the benchmark's own verdict — plus whether the run's own outcome was success. The map's caption names the grader and its weights.
- Denominator: Graded runs — runs the study's grader scored.
- Pairwise verdict
- A blind, position-debiased comparison of the two arms' final answers. It sees answer text only — cost and latency are measured separately.
- Denominator: Task-cell pairs where both arms produced a response.
- Infra-excluded
- A run killed by our own infrastructure. It is a missing measurement, never a loss for the arm, and is excluded from every rate denominator.
- Denominator: Reported as a count beside every affected panel.
- A/B arms
- Every task runs twice per harness × model cell — once without Web Research (baseline), once with it — on the same prompt, same model, cold start for both arms.
- Pre-registered rubrics
- Each task's acceptance criteria (5 per task) are written at task generation, before any run exists, so grading can never be shaped by the results.
- Judge panel
- A panel of independent judges (claude-opus-5, gpt-5.5, claude-fable-5), each at provider-default sampling (claude-opus-5, gpt-5.5 and claude-fable-5 do not accept a temperature setting), scores every successful response against its rubric; a criterion passes only when a strict majority of the panel passes it. Each judge sees only the task, the rubric, and the response — never which arm produced it, never token counts.
- Blind pairwise
- Each task's two responses are also compared blind as “Response A” and “Response B” by every panel judge, each judging twice with the order swapped (disagreement = tie); the pair's verdict is the panel's strict majority — no majority counts as a tie.
- How the head-to-head is decided
- The conclusion is not the blind judge's alone — that judge sees only the two response texts, never the rubric score or what each run cost. Evidence is ranked in order: whether the task's goal was achieved, then the rubric score, then the measured cost of getting there (tokens, turns, duration, spend), and only then the blind judge, for pairs nothing else separates. Cost never outranks correctness, and when neither arm achieved the goal the efficiency gap between them decides nothing.
- One run per task
- Every task ran once per cell and arm, in one phrasing. Run-to-run variation is therefore not measured: a single task's difference can be one lucky or unlucky attempt. Every row of the reference table treats the TASKS as the unit: its intervals resample the tasks and its p-values come from a paired test over them, so they describe how much the result depends on which tasks were drawn — not how it would change if the same tasks were run again. With few tasks no difference can reach significance (five tasks cannot go below p = 1/16).
- Sample size
- 2 cells × 5 tasks × 2 arms = 20 runs.