Do agent upgrades deliver what they promise? We measure.
Each study takes one subject's promise — a skill, a tool policy, an add-on — and measures it across real harness × model cells: A/B studies run every task under both arms and read the promise off the deltas; shootouts run one arm per cell and rank the field.
Kitesurf (DevTools)
50 read-only intent tasks from benchme 0.6.0's intent corpus (warehouse 26, helpdesk 13, vaultdocs 11), drawn per app in proportion to the qualifying tasks with the seed 'kitesurf-ab-benchme-50'. The site is the public hosted benchme at https://benchme.agentfront.sh. Every run first creates its own anonymous workspace (POST /api/workspaces with scenario acme-v1, seed 4242), so no two runs share…
2 cells · 50 tasks1views
Kitesurf (Playwright)
50 read-only intent tasks from benchme 0.6.0's intent corpus (warehouse 26, helpdesk 13, vaultdocs 11), drawn per app in proportion to the qualifying tasks with the seed 'kitesurf-ab-benchme-50'. The site is the public hosted benchme at https://benchme.agentfront.sh. Every run first creates its own anonymous workspace (POST /api/workspaces with scenario acme-v1, seed 4242), so no two runs share…
2 cells · 50 tasks2views
WindTunnel
WindTunnel, by nekuda, measures how well AI agents complete real tasks on the web. Across eight realistic sites (shops, booking, ticketing, a CRM, a learning platform, a blog, a directory and a members-only app), agents look up answers or take actions, and every task is scored automatically against the site itself.
2 cells · 50 tasks4views
WindTunnel
WindTunnel, by nekuda, measures how well AI agents complete real tasks on the web. Across eight realistic sites (shops, booking, ticketing, a CRM, a learning platform, a blog, a directory and a members-only app), agents look up answers or take actions, and every task is scored automatically against the site itself.
2 cells · 15 tasks5views
WindTunnel
WindTunnel, by nekuda, measures how well AI agents complete real tasks on the web. Across eight realistic sites (shops, booking, ticketing, a CRM, a learning platform, a blog, a directory and a members-only app), agents look up answers or take actions, and every task is scored automatically against the site itself.
Top of the board claude-code / claude-haiku-4-5-20251001
2 cells · 50 tasks8views
Web Research
we research skill improves searches
2 cells · 5 tasks1views
benchme — intent tasks
benchme, by orabenchmarks, measures how well AI agents handle everyday work in a company's internal web apps: a warehouse, a helpdesk, a document vault and a public website. Each task is a plain-language request, like placing an order or resolving a ticket, scored automatically against the result in the app.
Top of the board claude-code / claude-haiku-4-5-20251001
1 cell · 20 tasks1views