Do agent upgrades deliver what they promise? We measure.

Each study takes one subject's promise — a skill, a tool policy, an add-on — and measures it across real harness × model cells: A/B studies run every task under both arms and read the promise off the deltas; shootouts run one arm per cell and rank the field.

Kitesurf (DevTools)

50 read-only intent tasks from benchme 0.6.0's intent corpus (warehouse 26, helpdesk 13, vaultdocs 11), drawn per app in proportion to the qualifying tasks with the seed 'kitesurf-ab-benchme-50'. The site is the public hosted benchme at https://benchme.agentfront.sh. Every run first creates its own anonymous workspace (POST /api/workspaces with scenario acme-v1, seed 4242), so no two runs share…

+3ppgoal achievement(+6% relative)
49%52%goal achievement
claude-code — claude-haiku-4-5-20251001codex — gpt-5.4-mini

2 cells · 50 tasks1views

PartialCustomSep 29, 2026

Kitesurf (Playwright)

50 read-only intent tasks from benchme 0.6.0's intent corpus (warehouse 26, helpdesk 13, vaultdocs 11), drawn per app in proportion to the qualifying tasks with the seed 'kitesurf-ab-benchme-50'. The site is the public hosted benchme at https://benchme.agentfront.sh. Every run first creates its own anonymous workspace (POST /api/workspaces with scenario acme-v1, seed 4242), so no two runs share…

+4ppgoal achievement(+8% relative)
48%52%goal achievement
claude-code — claude-haiku-4-5-20251001codex — gpt-5.4-mini

2 cells · 50 tasks2views

PartialCustomSep 29, 2026

WindTunnel

WindTunnel, by nekuda, measures how well AI agents complete real tasks on the web. Across eight realistic sites (shops, booking, ticketing, a CRM, a learning platform, a blog, a directory and a members-only app), agents look up answers or take actions, and every task is scored automatically against the site itself.

+0ppcompletion rate(+0% relative)
100%100%completion rate
claude-code — claude-haiku-4-5-20251001codex — gpt-5.4-mini

2 cells · 50 tasks4views

DeliveredBenchmarkSep 28, 2026

WindTunnel

WindTunnel, by nekuda, measures how well AI agents complete real tasks on the web. Across eight realistic sites (shops, booking, ticketing, a CRM, a learning platform, a blog, a directory and a members-only app), agents look up answers or take actions, and every task is scored automatically against the site itself.

+3ppcompletion rate(+3% relative)
97%100%completion rate
claude-code — claude-haiku-4-5-20251001codex — gpt-5.4-mini

2 cells · 15 tasks5views

PartialBenchmarkSep 28, 2026

WindTunnel

WindTunnel, by nekuda, measures how well AI agents complete real tasks on the web. Across eight realistic sites (shops, booking, ticketing, a CRM, a learning platform, a blog, a directory and a members-only app), agents look up answers or take actions, and every task is scored automatically against the site itself.

Top of the board claude-code / claude-haiku-4-5-20251001

claude-code — claude-haiku-4-5-20251001codex — gpt-5.4-mini

2 cells · 50 tasks8views

MeasuredBenchmarkSep 28, 2026

Web Research

we research skill improves searches

+39%duration
50%50%goal achievement
claude-code — claude-haiku-4-5-20251001codex — gpt-5.4-mini

2 cells · 5 tasks1views

Not deliveredSkillSep 28, 2026

benchme — intent tasks

benchme, by orabenchmarks, measures how well AI agents handle everyday work in a company's internal web apps: a warehouse, a helpdesk, a document vault and a public website. Each task is a plain-language request, like placing an order or resolving a ticket, scored automatically against the result in the app.

Top of the board claude-code / claude-haiku-4-5-20251001

claude-code — claude-haiku-4-5-20251001

1 cell · 20 tasks1views

MeasuredBenchmarkSep 28, 2026