Partial · Study2views
benchme — in-pod Playwright vs Kitesurf (50 tasks)
Kitesurf (Playwright) vs Playwright (in-pod) · 50 tasks · 2 harnesses · 2 models
Playwright (in-pod) is the better choice over Kitesurf (Playwright) — cost +18%, tokens +27.6%, quality +1.7pts, with the goal gain concentrated in one cell rather than uniform.
Correctness did not separate the arms beyond the decision margin, and the baseline got there with 54% faster and 22% fewer tokens.
Abstract
Across 50 tasks over 2 harness/model cells, Kitesurf (Playwright) ran more expensively than Playwright (in-pod) — cost +18%, a median of $0.05 against $0.04 (means $0.05 against $0.05); tokens +27.6%, wall-clock +118.5%. The tail runs the other way: the treatment's worst case is worse than the baseline's. Quality improved: rubric score 88/100 against 87/100 (+1.7pts), and blind pairwise judging found 91% equivalent, 4% treatment better, 5% baseline better (of 100 judged pairs). Deterministic checks passed more often: 98% against 96% (+2pp (+2.1% relative)). Goal achievement moved 48% to 52%, +4pp (+8.3% relative). That gain is concentrated rather than general — codex/gpt-5.4-mini drives it (+4pp pooled; per-cell +2pp to +6pp). Harness and model are not crossed in this matrix, so neither can be credited alone.
The result
Playwright (in-pod)Kitesurf (Playwright)
Best arm overall: Playwright (in-pod) (72.3 vs 68.4 of 100)
Best setup: claude-code / claude-haiku-4-5-20251001 · Playwright (in-pod) (84.1 of 100)
- 1claude-code / claude-haiku-4-5-20251001Playwright (in-pod)84.1best
- 2claude-code / claude-haiku-4-5-20251001Kitesurf (Playwright)79.8
- 3codex / gpt-5.4-miniPlaywright (in-pod)65.0
- 4codex / gpt-5.4-miniKitesurf (Playwright)62.0
Overall score per arm, 0–100 points (not a pass rate): 30% goal + 30% quality + 15% evals + 25% efficiency (cost · tokens · duration, vs the board's best arm). Absent components renormalize. Best at the top — the board's best setup is tagged.
Head-to-head
Correctness did not separate the arms beyond the decision margin, and the baseline got there with 54% faster and 22% fewer tokens.
Same result both sides — decided on what it cost
What the blind judge alone said
Playwright (in-pod) wins 5Ties 91Treatment wins 4Win rate 4%
In blind position-debiased comparison, Kitesurf (Playwright) won 4 of 100 pairs against Playwright (in-pod); 91 were ties.
Quality × efficiency clusters — normalized per task
Every graded run, standardized WITHIN its task so difficulty cancels out: → right = fewer tokens than the field on the same task, ↑ up = higher quality index (70% rubric pass rate + 15% intent + 15% outcome) than the field. The field pools every harness and model, so a setup's left–right position largely reflects its own token habits. Dots are runs; each harness logo is a harness × model × arm centroid — hover it for the model and averages. Up-right wins.
4 runs not plotted (missing a grade or token count — infra failures included).
Every metric, per harness × model
Task completion
- claude-code / claude-haiku-4-5-20251001100% → 100%even
- codex / gpt-5.4-mini100% → 100%even
Goal achievement
- claude-code / claude-haiku-4-5-2025100188% → 90%2.0 pp better
- codex / gpt-5.4-mini8% → 14%6.0 pp better
Rubric quality /100
- claude-code / claude-haiku-4-5-2025100196 → 982.1 pts better
- codex / gpt-5.4-mini77 → 79even
Evals pass rate
- claude-code / claude-haiku-4-5-2025100196% → 98%2.0 pp better
- codex / gpt-5.4-mini96% → 98%2.0 pp better
Cost per run
- codex / gpt-5.4-mini$0.03 → $0.038% worse
- claude-code / claude-haiku-4-5-20251001$0.05 → $0.0626% worse
Tokens per run
- codex / gpt-5.4-mini88.2k → 102k16% worse
- claude-code / claude-haiku-4-5-20251001274.4k → 376k37% worse
Duration per run
- codex / gpt-5.4-mini30.9s → 55.8s80% worse
- claude-code / claude-haiku-4-5-2025100136.5s → 1m 27s139% worse
Rubric criteria passed
- claude-code / claude-haiku-4-5-2025100196% → 98%2.1 pp better
- codex / gpt-5.4-mini77% → 78%even
Playwright (in-pod)
Kitesurf (Playwright)
Deterministic checks
Pre-registered pass/fail expectations — MCP calls, commands, files, responses — checked deterministically against each run's recorded evidence, independent of any LLM judge.
Pass rate — Playwright (in-pod) vs Kitesurf (Playwright)96% vs 98% · +2ppShow per task (50)Hide
Distributions
Every completed run is one dot — the spread the averages hide. Click a dot to replay that run's journey.
Rubric quality per run
Token composition — average per run
Where agents fail
Each cell is the pass fraction of one pre-registered criterion — darker red, more failures. Click a fraction to read the judge's notes on the runs behind it.
| Criterion | claude-code | codex | All |
|---|---|---|---|
| T2The response text itself (not the run trace) identifies the product record the price comes from, e.g. by the product name and an identifier the warehouse app shows, so a reader can tell which record the figure relies on. | 0/2 | ||
| T9The response text itself (not the run trace) identifies the product record the price comes from, e.g. by the product name and an identifier the warehouse app shows, so a reader can tell which record the figure relies on. | 0/2 | ||
| T26The response states the basis for its pick: the unit price it found for the chosen product and the prices it compared it against, rather than a bare SKU with no stated basis. | 0/2 | ||
| T27The "city" value in the final JSON object is the string "Tessalind" (case-insensitive, surrounding whitespace ignored). | 0/2 | ||
| T49The response shows its basis: it names the document it read the figure from and identifies it as the final FY2025 results summary, not the superseded draft the task warns about. | 0/2 | ||
| T1The response text itself (not the run trace) names the document the id belongs to, e.g. its title as the vault shows it, so a reader can see which record the answer relies on. | 1/2 | ||
| T2The response text presents the figure as the unit price of the Quenquen Renmar kit, the product the task names, and not as another product's price or a total for several units. | 1/2 | ||
| T3The response answers the request as asked: the value it reports is a customer code (not only a customer name or another identifier), and the response text presents that customer as gold-tier and based in Quillhaven, addressing each condition the task set. | 1/2 | ||
| T3The response text itself (not the run trace) identifies the customer record the code came from, e.g. the customer's name as the warehouse app shows it, so a reader can see which record the answer relies on. | 1/2 | ||
| T4The response text presents the figure as the current in-stock quantity of the Falnexmar Kael assembly (SKU ENC-1107) across all depots, as asked, and not as the count at an individual depot or a figure for another item. | 1/2 | ||
| T4The response text itself (not the run trace) says which item record in the warehouse app the figure comes from, so a reader can trace where the number came from. | 1/2 | ||
| T5The response text itself (not the run trace) says which item record in the warehouse app the figure comes from, so a reader can trace where the number came from. | 1/2 | ||
| T6The response text presents the figure as the current in-stock quantity of the Zyom Renom kit (SKU FAS-1000) across all depots, as asked, and not as the count at an individual depot or a figure for another item. | 1/2 | ||
| T6The response text itself (not the run trace) says which item record in the warehouse app the figure comes from, so a reader can trace where the number came from. | 1/2 | ||
| T7The response text itself (not the run trace) identifies the item and the depot entry the figure was read from, so a reader can trace where the number came from. | 1/2 | ||
| T8The response text itself (not the run trace) identifies the item and the depot entry the figure was read from, so a reader can trace where the number came from. | 1/2 | ||
| T9The response text presents the figure as the unit price of the Lolo Vivi pack, the product the task names, and not as another product's price or a total for several units. | 1/2 | ||
| T10The response text presents the figure as the unit price of the Vizynex Renquen pack, the product the task names, and not as another product's price or a total for several units. | 1/2 | ||
| T10The response text itself (not the run trace) identifies the product record the price comes from, e.g. by the product name and an identifier the warehouse app shows, so a reader can tell which record the figure relies on. | 1/2 | ||
| T11The response shows its basis: it names the customer the code belongs to and states that this customer is platinum-tier and based in Karrow, addressing the conditions the task set. | 1/2 | ||
| T12The response names the customer the task asked about (Quenfalsil Industrial) together with the tier, so the value is clearly tied to that customer and not another. | 1/2 | ||
| T13The response names the customer the task asked about (Kasilvi Systems) together with the tier, so the value is clearly tied to that customer and not another. | 1/2 | ||
| T14The response makes clear the count is of the open orders of Vitru Industrial (C-127) and shows its basis: it identifies the orders it counted (e.g. by order number) or says how it decided which of that customer's orders were open. | 1/2 | ||
| T15The response makes clear the count is of the open orders of Thomarom Systems (C-115) and shows its basis: it identifies the orders it counted (e.g. by order number) or says how it decided which of that customer's orders were open. | 1/2 | ||
| T16The response names the order it looked up (SO-20003) alongside the status, so the value is clearly tied to that order and not another. | 1/2 | ||
| T17The response names the order it looked up (SO-20040) alongside the status, so the value is clearly tied to that order and not another. | 1/2 | ||
| T18The response names the order it looked up (SO-20167) alongside the status, so the value is clearly tied to that order and not another. | 1/2 | ||
| T19The response shows the basis for the total: it names the order it read (SO-20296) and either lists the line figures it summed or says which of that order's own figures it relied on. | 1/2 | ||
| T20The response shows the basis for the total: it names the order it read (SO-20102) and either lists the line figures it summed or says which of that order's own figures it relied on. | 1/2 | ||
| T21The response identifies its value as the total of order SO-20156 across all of that order's lines (it names that order), not as a single line's amount or another order's figure. | 1/2 | ||
| T22The response identifies its value as the quantity of SKU ENC-1075 on order SO-20060 (it names that SKU or order), not as a different figure such as the order's total units or stock on hand. | 1/2 | ||
| T23The response identifies its value as the quantity of SKU CAB-1034 on order SO-20118 (it names that SKU or order), not as a different figure such as the order's total units or stock on hand. | 1/2 | ||
| T24The response identifies the code as the destination of transfer TR-532 (it names that transfer), not as that transfer's origin or another transfer's destination. | 1/2 | ||
| T25The response frames its answer as the depot with the most units of SKU FLU-1055 specifically (it names that SKU or product), not as the depot with the most stock overall. | 1/2 | ||
| T25The response states the basis for its choice: the stock figures for SKU FLU-1055 it compared across depots, rather than a bare depot code with no stated basis. | 1/2 | ||
| T26The response frames its pick as the enclosures-category product with the highest unit price (it states unit price as the ranking criterion), not as a pick by stock level, order volume or popularity. | 1/2 | ||
| T27The response identifies the city as the location of depot QHV (it names that depot by code or name), not as the city of a different depot. | 1/2 | ||
| T31The response presents the value as the priority of ticket HD-5128 itself (it names that ticket), not the ticket's status or the priority of a different ticket. | 1/2 | ||
| T32The response presents the value as the priority of ticket HD-5136 itself (it names that ticket), not the ticket's status or the priority of a different ticket. | 1/2 | ||
| T35The response presents the agent as the assignee of ticket HD-5025 (it names that ticket), not the ticket's requester (the customer) or the agent on a different ticket. | 1/2 | ||
| T38The response presents the value as the team of support agent Lior Kowalczyk (it names that agent), not a different agent's team. | 1/2 | ||
| T39The response presents the value as the team of support agent Quill Ferrante (it names that agent), not a different agent's team. | 1/2 | ||
| T40The response presents the number as the count of helpdesk tickets reported by customer Omthonex Supply (C-131) (it names that customer), not tickets handled by an agent or another customer's tickets; if it counted only some of that customer's tickets, it says so. | 1/2 | ||
| T40The response states the basis for its count: it identifies the tickets it counted (e.g. by ticket number) or the record or listing it took the figure from, rather than giving a bare number. | 1/2 | ||
| T41The response ties the id to the document the task asked about: alongside the id it names that document (e.g. its title as shown in the vault), and the document it names is operations memo 2: rentruzy retrospective, not a different memo. | 1/2 | ||
| T42The response labels the value it reports as the standard supplier lead time for the Zyom Renom kit (SKU FAS-1000) — it names that product or SKU — not as a different kind of figure or another product's figure. | 1/2 | ||
| T42The response shows its basis: it names the vault document it read the figure from (e.g. its title as shown in the vault), so a reader can see that the answer relies on the datasheet for SKU FAS-1000. | 1/2 | ||
| T43The response labels the value it reports as the standard supplier lead time for the Lofalel Zynex assembly (SKU ADH-1054) — it names that product or SKU — not as a different kind of figure or another product's figure. | 1/2 | ||
| T43The response shows its basis: it names the vault document it read the figure from (e.g. its title as shown in the vault), so a reader can see that the answer relies on the datasheet for SKU ADH-1054. | 1/2 | ||
| T44The response labels the value it reports as the minimum order quantity for the Baelfal Marren pack (SKU CAB-1018) — it names that product or SKU — not as a different kind of figure or another product's figure. | 1/2 | ||
| T44The response shows its basis: it names the vault document it read the figure from (e.g. its title as shown in the vault), so a reader can see that the answer relies on the datasheet for SKU CAB-1018. | 1/2 | ||
| T45The response labels the value it reports as the discount off list price for platinum customers under the customer tier discounts policy (it names the platinum tier), not as another tier's discount or a different term of that policy. | 1/2 | ||
| T45The response shows its basis: it names the vault document it read the figure from (e.g. its title as shown in the vault), so a reader can see that the answer relies on the customer tier discounts policy. | 1/2 | ||
| T46The response labels the value it reports as the number of days after delivery within which customers may return unused goods, not as a time limit for something else or a figure in other units. | 1/2 | ||
| T46The response shows its basis: it names the vault document it read the figure from (e.g. its title as shown in the vault), so a reader can see that the answer relies on the returns and warranty policy. | 1/2 | ||
| T47The response labels the hour it reports as the time until which Thoomsil depot 3 (QHV) dispatches goods — it names that depot by name or code — not as some other hour of that depot or another depot's hour. | 1/2 | ||
| T47The response shows its basis: it names the vault document it read the figure from (e.g. its title as shown in the vault), so a reader can see that the answer relies on the depot operating hours policy. | 1/2 | ||
| T48The response labels the value it reports as the number of months until Trumar Systems' next review — it names that account — not as a calendar date, a duration in other units, or another account's review timing. | 1/2 | ||
| T48The response shows its basis: it names the vault document it read the figure from (e.g. its title as shown in the vault), so a reader can see that the answer relies on the account review notes for Trumar Systems. | 1/2 | ||
| T49The response labels the value it reports as the company's number of employees for FY2025, not as a different kind of figure or a figure for another period. | 1/2 | ||
| T50The response labels the value it reports as the company's FY2025 revenue in US dollars, not as a different kind of figure, a figure in another currency, or a figure for another period. | 1/2 | ||
| T50The response shows its basis: it names the document it read the figure from and identifies it as the final FY2025 results summary, not the superseded draft the task warns about. | 1/2 |
Run outcomes
How each run ended — whether the agent finished, not whether its answer passed. Whether answers passed is in the results above.
Efficiency frontier
Statistics
| Metric | Arm | n | Mean | Median (pooled) | Min | Max | Std dev |
|---|---|---|---|---|---|---|---|
| Rubric quality | base | 100/100 | 87 | 83 | 17 | 100 | 15 |
| treat | 100/100 | 88 | 100 | 60 | 100 | 13 | |
| Cost | base | 98/100 | $0.05 | $0.03 | $0.02 | $0.30 | $0.04 |
| treat | 98/100 | $0.05 | $0.05 | $0.02 | $0.37 | $0.04 | |
| Duration | base | 100/100 | 35.9s | 32.8s | 24.4s | 1m 47s | 11.3s |
| treat | 100/100 | 1m 19s | 1m 13s | 44.8s | 4m 09s | 31.6s | |
| Total tokens | base | 98/100 | 212.6k | 157.6k | 62.2k | 1.8M | 230.6k |
| treat | 98/100 | 270.4k | 240k | 77.5k | 2.1M | 250k | |
| Output tokens | base | 98/100 | 1.1k | 932 | 452 | 5.5k | 772 |
| treat | 98/100 | 1.3k | 1.1k | 560 | 6.8k | 835 | |
| Cache-read tokens | base | 98/100 | 195.3k | 149.4k | 47.6k | 1.7M | 222.1k |
| treat | 98/100 | 253.3k | 229.6k | 59.8k | 2M | 240k | |
| Cache-write tokens | base | 98/100 | 9.5k | 0 | 0 | 81.7k | 14k |
| treat | 98/100 | 9k | 3.8k | 0 | 105.2k | 13.9k | |
| Turns | base | 100/100 | 12 | 9.5 | 3 | 74 | 11.4 |
| treat | 100/100 | 15.2 | 10 | 3 | 83 | 13.5 |
n = runs carrying the fact / completed runs in the arm; every statistic runs over present facts only — a missing fact is never counted as 0. Std dev is the sample form (n−1), withheld below n = 2. Every figure here POOLS all runs. The typical-run figures in the abstract, the head-to-head decision and the per-setup panels take each task's median first and then the median across tasks, so a task that ran more often never outweighs one that ran once; the two medians can differ. A promise study's headline (“cut average tokens”) compares per-arm means.
Tasks
First create your own workspace: POST https://benchme.agentfront.sh/api/workspaces with the JSON body {"scenario":"acme-v1","seed":4242}. That is the only request you may make outside the browser. Open the vaultdocs URL it returns (urls.apps.vaultdocs) in the browser and work only there — do not call the app's REST API, its /mcp endpoint or its /ask endpoint. Search the document vault for the document about returns and warranty policy and report its document id as JSON: {"documentId": "<id>"}. Give that JSON object in your final reply.
benchme intent-doc-01
Pre-registered rubric
- The "documentId" value in the final JSON object is the string "doc-101" (case-insensitive, surrounding whitespace ignored).
- The final reply contains a JSON object with the key "documentId" whose value is a quoted JSON string, as the task requested.
- The response answers the request as asked: the value it reports is a document id (not only a title, URL or search snippet), and the response text presents it as the id of the returns and warranty policy document.
- The response text itself (not the run trace) names the document the id belongs to, e.g. its title as the vault shows it, so a reader can see which record the answer relies on.
- The response commits to its answer: it gives no alternative or conflicting "documentId" values, and if it mentions other candidate documents, it says why it did not use them.
- The response is candid about gaps: if it says a step failed or was incomplete (creating the workspace, opening the vaultdocs app, the search), it does not also present a "documentId" value as found, and it never gives a placeholder such as "<id>" as the value.
First create your own workspace: POST https://benchme.agentfront.sh/api/workspaces with the JSON body {"scenario":"acme-v1","seed":4242}. That is the only request you may make outside the browser. Open the warehouse URL it returns (urls.apps.warehouse) in the browser and work only there — do not call the app's REST API, its /mcp endpoint or its /ask endpoint. What is the unit price of the Quenquen Renmar kit? Answer as JSON: {"price": <number, USD>}. Give that JSON object in your final reply.
benchme intent-price-01
Pre-registered rubric
- The "price" value in the final JSON object is the number 50.25, within ±0.005 (a JSON number, not a quoted string).
- The final reply contains a JSON object with the key "price" whose value is a bare JSON number (not a quoted string, and with no currency symbol or unit text inside the value).
- The response text presents the figure as the unit price of the Quenquen Renmar kit, the product the task names, and not as another product's price or a total for several units.
- The response text itself (not the run trace) identifies the product record the price comes from, e.g. by the product name and an identifier the warehouse app shows, so a reader can tell which record the figure relies on.
- The response commits to its answer: it gives no alternative or conflicting "price" values, and if it mentions other candidate products or prices, it says why it did not use them.
- The response is candid about gaps: if it says a step failed or was incomplete (creating the workspace, opening the warehouse app, finding the product), it does not also present a "price" value as found, and it never gives a placeholder such as "<number, USD>" as the value.
First create your own workspace: POST https://benchme.agentfront.sh/api/workspaces with the JSON body {"scenario":"acme-v1","seed":4242}. That is the only request you may make outside the browser. Open the warehouse URL it returns (urls.apps.warehouse) in the browser and work only there — do not call the app's REST API, its /mcp endpoint or its /ask endpoint. Which of our gold-tier customers is based in Quillhaven? Answer as JSON: {"customerCode": "<code>"}. Give that JSON object in your final reply.
benchme intent-customer-01
Pre-registered rubric
- The "customerCode" value in the final JSON object is the string "C-114" (case-insensitive, surrounding whitespace ignored).
- The final reply contains a JSON object with the key "customerCode" whose value is a quoted JSON string, as the task requested.
- The response answers the request as asked: the value it reports is a customer code (not only a customer name or another identifier), and the response text presents that customer as gold-tier and based in Quillhaven, addressing each condition the task set.
- The response text itself (not the run trace) identifies the customer record the code came from, e.g. the customer's name as the warehouse app shows it, so a reader can see which record the answer relies on.
- The response commits to its answer: it gives no alternative or conflicting "customerCode" values, and if it mentions other candidate customers, it says why it did not use them.
- The response is candid about gaps: if it says a step failed or was incomplete (creating the workspace, opening the warehouse app, finding the customer), it does not also present a "customerCode" value as found, and it never gives a placeholder such as "<code>" as the value.
First create your own workspace: POST https://benchme.agentfront.sh/api/workspaces with the JSON body {"scenario":"acme-v1","seed":4242}. That is the only request you may make outside the browser. Open the warehouse URL it returns (urls.apps.warehouse) in the browser and work only there — do not call the app's REST API, its /mcp endpoint or its /ask endpoint. How many units of the Falnexmar Kael assembly (SKU ENC-1107) are currently in stock across all depots? Answer as JSON: {"qty": <number>}. Give that JSON object in your final reply.
benchme intent-stock-02
Pre-registered rubric
- The "qty" value in the final JSON object is the number 263 exactly (a JSON number, not a quoted string).
- The final reply contains a JSON object with the key "qty" whose value is a bare JSON number (not a quoted string, and with no unit text inside the value).
- The response text presents the figure as the current in-stock quantity of the Falnexmar Kael assembly (SKU ENC-1107) across all depots, as asked, and not as the count at an individual depot or a figure for another item.
- The response text itself (not the run trace) says which item record in the warehouse app the figure comes from, so a reader can trace where the number came from.
- The response commits to its answer: it gives no alternative or conflicting "qty" values, and if it mentions other candidate figures, it says why it did not use them.
- The response is candid about gaps: if it says a step failed or was incomplete (creating the workspace, opening the warehouse app, finding the item, covering the full scope asked), it does not also present a "qty" value as found, and it never gives a placeholder such as "<number>" as the value.
First create your own workspace: POST https://benchme.agentfront.sh/api/workspaces with the JSON body {"scenario":"acme-v1","seed":4242}. That is the only request you may make outside the browser. Open the warehouse URL it returns (urls.apps.warehouse) in the browser and work only there — do not call the app's REST API, its /mcp endpoint or its /ask endpoint. How many units of the Quenel Omren module (SKU TOO-1085) are currently in stock across all depots? Answer as JSON: {"qty": <number>}. Give that JSON object in your final reply.
benchme intent-stock-04
Pre-registered rubric
- The "qty" value in the final JSON object is the number 508 exactly (a JSON number, not a quoted string).
- The final reply contains a JSON object with the key "qty" whose value is a bare JSON number (not a quoted string, and with no unit text inside the value).
- The response text presents the figure as the current in-stock quantity of the Quenel Omren module (SKU TOO-1085) across all depots, as asked, and not as the count at an individual depot or a figure for another item.
- The response text itself (not the run trace) says which item record in the warehouse app the figure comes from, so a reader can trace where the number came from.
- The response commits to its answer: it gives no alternative or conflicting "qty" values, and if it mentions other candidate figures, it says why it did not use them.
- The response is candid about gaps: if it says a step failed or was incomplete (creating the workspace, opening the warehouse app, finding the item, covering the full scope asked), it does not also present a "qty" value as found, and it never gives a placeholder such as "<number>" as the value.
First create your own workspace: POST https://benchme.agentfront.sh/api/workspaces with the JSON body {"scenario":"acme-v1","seed":4242}. That is the only request you may make outside the browser. Open the warehouse URL it returns (urls.apps.warehouse) in the browser and work only there — do not call the app's REST API, its /mcp endpoint or its /ask endpoint. How many units of the Zyom Renom kit (SKU FAS-1000) are currently in stock across all depots? Answer as JSON: {"qty": <number>}. Give that JSON object in your final reply.
benchme intent-stock-05
Pre-registered rubric
- The "qty" value in the final JSON object is the number 912 exactly (a JSON number, not a quoted string).
- The final reply contains a JSON object with the key "qty" whose value is a bare JSON number (not a quoted string, and with no unit text inside the value).
- The response text presents the figure as the current in-stock quantity of the Zyom Renom kit (SKU FAS-1000) across all depots, as asked, and not as the count at an individual depot or a figure for another item.
- The response text itself (not the run trace) says which item record in the warehouse app the figure comes from, so a reader can trace where the number came from.
- The response commits to its answer: it gives no alternative or conflicting "qty" values, and if it mentions other candidate figures, it says why it did not use them.
- The response is candid about gaps: if it says a step failed or was incomplete (creating the workspace, opening the warehouse app, finding the item, covering the full scope asked), it does not also present a "qty" value as found, and it never gives a placeholder such as "<number>" as the value.
First create your own workspace: POST https://benchme.agentfront.sh/api/workspaces with the JSON body {"scenario":"acme-v1","seed":4242}. That is the only request you may make outside the browser. Open the warehouse URL it returns (urls.apps.warehouse) in the browser and work only there — do not call the app's REST API, its /mcp endpoint or its /ask endpoint. How many units of the Loom Loquen kit (SKU OPT-1097) are held at depot OST (Zyvi depot 2)? Answer as JSON: {"qty": <number>}. Give that JSON object in your final reply.
benchme intent-depot-stock-03
Pre-registered rubric
- The "qty" value in the final JSON object is the number 371 exactly (a JSON number, not a quoted string).
- The final reply contains a JSON object with the key "qty" whose value is a bare JSON number (not a quoted string, and with no unit text inside the value).
- The response text presents the figure as the quantity of the Loom Loquen kit (SKU OPT-1097) held at depot OST (Zyvi depot 2), as asked, and not as a total across depots or a figure for another depot or item.
- The response text itself (not the run trace) identifies the item and the depot entry the figure was read from, so a reader can trace where the number came from.
- The response commits to its answer: it gives no alternative or conflicting "qty" values, and if it mentions other candidate figures, it says why it did not use them.
- The response is candid about gaps: if it says a step failed or was incomplete (creating the workspace, opening the warehouse app, finding the item at that depot), it does not also present a "qty" value as found, and it never gives a placeholder such as "<number>" as the value.
First create your own workspace: POST https://benchme.agentfront.sh/api/workspaces with the JSON body {"scenario":"acme-v1","seed":4242}. That is the only request you may make outside the browser. Open the warehouse URL it returns (urls.apps.warehouse) in the browser and work only there — do not call the app's REST API, its /mcp endpoint or its /ask endpoint. How many units of the Loloom Falvi unit (SKU FAS-1040) are held at depot QHV (Thoomsil depot 3)? Answer as JSON: {"qty": <number>}. Give that JSON object in your final reply.
benchme intent-depot-stock-04
Pre-registered rubric
- The "qty" value in the final JSON object is the number 89 exactly (a JSON number, not a quoted string).
- The final reply contains a JSON object with the key "qty" whose value is a bare JSON number (not a quoted string, and with no unit text inside the value).
- The response text presents the figure as the quantity of the Loloom Falvi unit (SKU FAS-1040) held at depot QHV (Thoomsil depot 3), as asked, and not as a total across depots or a figure for another depot or item.
- The response text itself (not the run trace) identifies the item and the depot entry the figure was read from, so a reader can trace where the number came from.
- The response commits to its answer: it gives no alternative or conflicting "qty" values, and if it mentions other candidate figures, it says why it did not use them.
- The response is candid about gaps: if it says a step failed or was incomplete (creating the workspace, opening the warehouse app, finding the item at that depot), it does not also present a "qty" value as found, and it never gives a placeholder such as "<number>" as the value.
First create your own workspace: POST https://benchme.agentfront.sh/api/workspaces with the JSON body {"scenario":"acme-v1","seed":4242}. That is the only request you may make outside the browser. Open the warehouse URL it returns (urls.apps.warehouse) in the browser and work only there — do not call the app's REST API, its /mcp endpoint or its /ask endpoint. What is the unit price of the Lolo Vivi pack? Answer as JSON: {"price": <number, USD>}. Give that JSON object in your final reply.
benchme intent-price-03
Pre-registered rubric
- The "price" value in the final JSON object is the number 23.5, within ±0.005 (a JSON number, not a quoted string).
- The final reply contains a JSON object with the key "price" whose value is a bare JSON number (not a quoted string, and with no currency symbol or unit text inside the value).
- The response text presents the figure as the unit price of the Lolo Vivi pack, the product the task names, and not as another product's price or a total for several units.
- The response text itself (not the run trace) identifies the product record the price comes from, e.g. by the product name and an identifier the warehouse app shows, so a reader can tell which record the figure relies on.
- The response commits to its answer: it gives no alternative or conflicting "price" values, and if it mentions other candidate products or prices, it says why it did not use them.
- The response is candid about gaps: if it says a step failed or was incomplete (creating the workspace, opening the warehouse app, finding the product), it does not also present a "price" value as found, and it never gives a placeholder such as "<number, USD>" as the value.
First create your own workspace: POST https://benchme.agentfront.sh/api/workspaces with the JSON body {"scenario":"acme-v1","seed":4242}. That is the only request you may make outside the browser. Open the warehouse URL it returns (urls.apps.warehouse) in the browser and work only there — do not call the app's REST API, its /mcp endpoint or its /ask endpoint. What is the unit price of the Vizynex Renquen pack? Answer as JSON: {"price": <number, USD>}. Give that JSON object in your final reply.
benchme intent-price-06
Pre-registered rubric
- The "price" value in the final JSON object is the number 148.25, within ±0.005 (a JSON number, not a quoted string).
- The final reply contains a JSON object with the key "price" whose value is a bare JSON number (not a quoted string, and with no currency symbol or unit text inside the value).
- The response text presents the figure as the unit price of the Vizynex Renquen pack, the product the task names, and not as another product's price or a total for several units.
- The response text itself (not the run trace) identifies the product record the price comes from, e.g. by the product name and an identifier the warehouse app shows, so a reader can tell which record the figure relies on.
- The response commits to its answer: it gives no alternative or conflicting "price" values, and if it mentions other candidate products or prices, it says why it did not use them.
- The response is candid about gaps: if it says a step failed or was incomplete (creating the workspace, opening the warehouse app, finding the product), it does not also present a "price" value as found, and it never gives a placeholder such as "<number, USD>" as the value.
First create your own workspace: POST https://benchme.agentfront.sh/api/workspaces with the JSON body {"scenario":"acme-v1","seed":4242}. That is the only request you may make outside the browser. Open the warehouse URL it returns (urls.apps.warehouse) in the browser and work only there — do not call the app's REST API, its /mcp endpoint or its /ask endpoint. Which of our platinum-tier customers is based in Karrow? Answer as JSON: {"customerCode": "<code>"}. Give that JSON object in your final reply.
benchme intent-customer-04
Pre-registered rubric
- The "customerCode" value in the final JSON object is the string "C-115" (case-insensitive, surrounding whitespace ignored).
- The final reply includes a JSON object of the shape the task requested: the key "customerCode" with its value, and no other keys in that object.
- The "customerCode" value is a customer code string — not a customer name, a sentence, a list of codes or the placeholder "<code>".
- The response shows its basis: it names the customer the code belongs to and states that this customer is platinum-tier and based in Karrow, addressing the conditions the task set.
- Any value for "customerCode" stated in the prose matches the value in the JSON object; the response gives no other, conflicting value.
- The response does not present a guessed, assumed or placeholder value as if it were read from the warehouse app: if a step failed or the record it needed could not be read, it says so explicitly and does not present the value as confirmed.
First create your own workspace: POST https://benchme.agentfront.sh/api/workspaces with the JSON body {"scenario":"acme-v1","seed":4242}. That is the only request you may make outside the browser. Open the warehouse URL it returns (urls.apps.warehouse) in the browser and work only there — do not call the app's REST API, its /mcp endpoint or its /ask endpoint. What tier is our customer Quenfalsil Industrial? Answer as JSON: {"tier": "<tier>"}. Give that JSON object in your final reply.
benchme intent-tier-01
Pre-registered rubric
- The "tier" value in the final JSON object is the string "gold" (case-insensitive, surrounding whitespace ignored).
- The final reply includes a JSON object of the shape the task requested: the key "tier" with its value, and no other keys in that object.
- The "tier" value is a tier label string on its own — not a sentence, a list of tiers or the placeholder "<tier>".
- The response names the customer the task asked about (Quenfalsil Industrial) together with the tier, so the value is clearly tied to that customer and not another.
- Any value for "tier" stated in the prose matches the value in the JSON object; the response gives no other, conflicting value.
- The response does not present a guessed, assumed or placeholder value as if it were read from the warehouse app: if a step failed or the record it needed could not be read, it says so explicitly and does not present the value as confirmed.
First create your own workspace: POST https://benchme.agentfront.sh/api/workspaces with the JSON body {"scenario":"acme-v1","seed":4242}. That is the only request you may make outside the browser. Open the warehouse URL it returns (urls.apps.warehouse) in the browser and work only there — do not call the app's REST API, its /mcp endpoint or its /ask endpoint. What tier is our customer Kasilvi Systems? Answer as JSON: {"tier": "<tier>"}. Give that JSON object in your final reply.
benchme intent-tier-04
Pre-registered rubric
- The "tier" value in the final JSON object is the string "gold" (case-insensitive, surrounding whitespace ignored).
- The final reply includes a JSON object of the shape the task requested: the key "tier" with its value, and no other keys in that object.
- The "tier" value is a tier label string on its own — not a sentence, a list of tiers or the placeholder "<tier>".
- The response names the customer the task asked about (Kasilvi Systems) together with the tier, so the value is clearly tied to that customer and not another.
- Any value for "tier" stated in the prose matches the value in the JSON object; the response gives no other, conflicting value.
- The response does not present a guessed, assumed or placeholder value as if it were read from the warehouse app: if a step failed or the record it needed could not be read, it says so explicitly and does not present the value as confirmed.
First create your own workspace: POST https://benchme.agentfront.sh/api/workspaces with the JSON body {"scenario":"acme-v1","seed":4242}. That is the only request you may make outside the browser. Open the warehouse URL it returns (urls.apps.warehouse) in the browser and work only there — do not call the app's REST API, its /mcp endpoint or its /ask endpoint. How many open orders does Vitru Industrial (C-127) currently have? Answer as JSON: {"count": <number>}. Give that JSON object in your final reply.
benchme intent-open-orders-03
Pre-registered rubric
- The "count" value in the final JSON object is the number 4 exactly (a JSON number, not a quoted string).
- The final reply includes a JSON object of the shape the task requested: the key "count" with its value, and no other keys in that object.
- The "count" value is a bare JSON number (a whole number) — not a quoted string, a range, a word or an expression.
- The response makes clear the count is of the open orders of Vitru Industrial (C-127) and shows its basis: it identifies the orders it counted (e.g. by order number) or says how it decided which of that customer's orders were open.
- Any value for "count" stated in the prose matches the value in the JSON object; the response gives no other, conflicting value.
- The response does not present a guessed, assumed or placeholder value as if it were read from the warehouse app: if a step failed or the record it needed could not be read, it says so explicitly and does not present the value as confirmed.
First create your own workspace: POST https://benchme.agentfront.sh/api/workspaces with the JSON body {"scenario":"acme-v1","seed":4242}. That is the only request you may make outside the browser. Open the warehouse URL it returns (urls.apps.warehouse) in the browser and work only there — do not call the app's REST API, its /mcp endpoint or its /ask endpoint. How many open orders does Thomarom Systems (C-115) currently have? Answer as JSON: {"count": <number>}. Give that JSON object in your final reply.
benchme intent-open-orders-04
Pre-registered rubric
- The "count" value in the final JSON object is the number 2 exactly (a JSON number, not a quoted string).
- The final reply includes a JSON object of the shape the task requested: the key "count" with its value, and no other keys in that object.
- The "count" value is a bare JSON number (a whole number) — not a quoted string, a range, a word or an expression.
- The response makes clear the count is of the open orders of Thomarom Systems (C-115) and shows its basis: it identifies the orders it counted (e.g. by order number) or says how it decided which of that customer's orders were open.
- Any value for "count" stated in the prose matches the value in the JSON object; the response gives no other, conflicting value.
- The response does not present a guessed, assumed or placeholder value as if it were read from the warehouse app: if a step failed or the record it needed could not be read, it says so explicitly and does not present the value as confirmed.
First create your own workspace: POST https://benchme.agentfront.sh/api/workspaces with the JSON body {"scenario":"acme-v1","seed":4242}. That is the only request you may make outside the browser. Open the warehouse URL it returns (urls.apps.warehouse) in the browser and work only there — do not call the app's REST API, its /mcp endpoint or its /ask endpoint. What is the status of order SO-20003? Answer as JSON: {"status": "<status>"}. Give that JSON object in your final reply.
benchme intent-order-status-01
Pre-registered rubric
- The "status" value in the final JSON object is the string "shipped" (case-insensitive, surrounding whitespace ignored).
- The final reply includes a JSON object of the shape the task requested: the key "status" with its value, and no other keys in that object.
- The "status" value is a status label string on its own — not a sentence, a list of statuses or the placeholder "<status>".
- The response names the order it looked up (SO-20003) alongside the status, so the value is clearly tied to that order and not another.
- Any value for "status" stated in the prose matches the value in the JSON object; the response gives no other, conflicting value.
- The response does not present a guessed, assumed or placeholder value as if it were read from the warehouse app: if a step failed or the record it needed could not be read, it says so explicitly and does not present the value as confirmed.
First create your own workspace: POST https://benchme.agentfront.sh/api/workspaces with the JSON body {"scenario":"acme-v1","seed":4242}. That is the only request you may make outside the browser. Open the warehouse URL it returns (urls.apps.warehouse) in the browser and work only there — do not call the app's REST API, its /mcp endpoint or its /ask endpoint. What is the status of order SO-20040? Answer as JSON: {"status": "<status>"}. Give that JSON object in your final reply.
benchme intent-order-status-02
Pre-registered rubric
- The "status" value in the final JSON object is the string "shipped" (case-insensitive, surrounding whitespace ignored).
- The final reply includes a JSON object of the shape the task requested: the key "status" with its value, and no other keys in that object.
- The "status" value is a status label string on its own — not a sentence, a list of statuses or the placeholder "<status>".
- The response names the order it looked up (SO-20040) alongside the status, so the value is clearly tied to that order and not another.
- Any value for "status" stated in the prose matches the value in the JSON object; the response gives no other, conflicting value.
- The response does not present a guessed, assumed or placeholder value as if it were read from the warehouse app: if a step failed or the record it needed could not be read, it says so explicitly and does not present the value as confirmed.
First create your own workspace: POST https://benchme.agentfront.sh/api/workspaces with the JSON body {"scenario":"acme-v1","seed":4242}. That is the only request you may make outside the browser. Open the warehouse URL it returns (urls.apps.warehouse) in the browser and work only there — do not call the app's REST API, its /mcp endpoint or its /ask endpoint. What is the status of order SO-20167? Answer as JSON: {"status": "<status>"}. Give that JSON object in your final reply.
benchme intent-order-status-04
Pre-registered rubric
- The "status" value in the final JSON object is the string "cancelled" (case-insensitive, surrounding whitespace ignored).
- The final reply includes a JSON object of the shape the task requested: the key "status" with its value, and no other keys in that object.
- The "status" value is a status label string on its own — not a sentence, a list of statuses or the placeholder "<status>".
- The response names the order it looked up (SO-20167) alongside the status, so the value is clearly tied to that order and not another.
- Any value for "status" stated in the prose matches the value in the JSON object; the response gives no other, conflicting value.
- The response does not present a guessed, assumed or placeholder value as if it were read from the warehouse app: if a step failed or the record it needed could not be read, it says so explicitly and does not present the value as confirmed.
First create your own workspace: POST https://benchme.agentfront.sh/api/workspaces with the JSON body {"scenario":"acme-v1","seed":4242}. That is the only request you may make outside the browser. Open the warehouse URL it returns (urls.apps.warehouse) in the browser and work only there — do not call the app's REST API, its /mcp endpoint or its /ask endpoint. What is the total value of order SO-20296 in US dollars — every line's quantity times its unit price, summed? Answer as JSON: {"total": <number>}. Give that JSON object in your final reply.
benchme intent-order-total-01
Pre-registered rubric
- The "total" value in the final JSON object is the number 7093.5, within ±0.005 (a JSON number, not a quoted string).
- The final reply includes a JSON object of the shape the task requested: the key "total" with its value, and no other keys in that object.
- The "total" value is a bare JSON number — not a quoted string, and with no "$", currency code, unit or thousands separator inside the JSON.
- The response shows the basis for the total: it names the order it read (SO-20296) and either lists the line figures it summed or says which of that order's own figures it relied on.
- Any total stated in the prose matches the JSON "total" (a "$" sign or thousands separators in the prose are fine); the response gives no other, conflicting total.
- The response does not present a guessed, assumed or placeholder value as if it were read from the warehouse app: if a step failed or the record it needed could not be read, it says so explicitly and does not present the value as confirmed.
First create your own workspace: POST https://benchme.agentfront.sh/api/workspaces with the JSON body {"scenario":"acme-v1","seed":4242}. That is the only request you may make outside the browser. Open the warehouse URL it returns (urls.apps.warehouse) in the browser and work only there — do not call the app's REST API, its /mcp endpoint or its /ask endpoint. What is the total value of order SO-20102 in US dollars — every line's quantity times its unit price, summed? Answer as JSON: {"total": <number>}. Give that JSON object in your final reply.
benchme intent-order-total-02
Pre-registered rubric
- The "total" value in the final JSON object is the number 8691.75, within ±0.005 (a JSON number, not a quoted string).
- The final reply includes a JSON object of the shape the task requested: the key "total" with its value, and no other keys in that object.
- The "total" value is a bare JSON number — not a quoted string, and with no "$", currency code, unit or thousands separator inside the JSON.
- The response shows the basis for the total: it names the order it read (SO-20102) and either lists the line figures it summed or says which of that order's own figures it relied on.
- Any total stated in the prose matches the JSON "total" (a "$" sign or thousands separators in the prose are fine); the response gives no other, conflicting total.
- The response does not present a guessed, assumed or placeholder value as if it were read from the warehouse app: if a step failed or the record it needed could not be read, it says so explicitly and does not present the value as confirmed.
First create your own workspace: POST https://benchme.agentfront.sh/api/workspaces with the JSON body {"scenario":"acme-v1","seed":4242}. That is the only request you may make outside the browser. Open the warehouse URL it returns (urls.apps.warehouse) in the browser and work only there — do not call the app's REST API, its /mcp endpoint or its /ask endpoint. What is the total value of order SO-20156 in US dollars — every line's quantity times its unit price, summed? Answer as JSON: {"total": <number>}. Give that JSON object in your final reply.
benchme intent-order-total-04
Pre-registered rubric
- The "total" value in the final JSON object is the number 16555.25, within ±0.005 (a JSON number, not a quoted string).
- The final reply contains a JSON object with the field "total" whose value is a single JSON number (not a quoted string, a range, or text with a currency symbol or thousands separators).
- The response identifies its value as the total of order SO-20156 across all of that order's lines (it names that order), not as a single line's amount or another order's figure.
- The response commits to a single value for "total"; it does not offer alternative, hedged or conflicting values.
- If the response shows supporting figures (line quantities, unit prices, line amounts or subtotals), they are arithmetically consistent with the "total" it reports.
- If the response reports that it could not read every line of order SO-20156, it marks its total as partial or unverified rather than presenting it as complete.
First create your own workspace: POST https://benchme.agentfront.sh/api/workspaces with the JSON body {"scenario":"acme-v1","seed":4242}. That is the only request you may make outside the browser. Open the warehouse URL it returns (urls.apps.warehouse) in the browser and work only there — do not call the app's REST API, its /mcp endpoint or its /ask endpoint. How many units of SKU ENC-1075 does order SO-20060 contain? Answer as JSON: {"qty": <number>}. Give that JSON object in your final reply.
benchme intent-order-line-01
Pre-registered rubric
- The "qty" value in the final JSON object is the number 9 exactly (a JSON number, not a quoted string).
- The final reply contains a JSON object with the field "qty" whose value is a single JSON number (not a quoted string, a range, or text such as "units").
- The response identifies its value as the quantity of SKU ENC-1075 on order SO-20060 (it names that SKU or order), not as a different figure such as the order's total units or stock on hand.
- The response commits to a single value for "qty"; it does not offer alternative, hedged or conflicting values.
- If the response reports any difficulty finding SKU ENC-1075 on order SO-20060 or reading that line, it marks its quantity as unverified rather than presenting it as confirmed.
First create your own workspace: POST https://benchme.agentfront.sh/api/workspaces with the JSON body {"scenario":"acme-v1","seed":4242}. That is the only request you may make outside the browser. Open the warehouse URL it returns (urls.apps.warehouse) in the browser and work only there — do not call the app's REST API, its /mcp endpoint or its /ask endpoint. How many units of SKU CAB-1034 does order SO-20118 contain? Answer as JSON: {"qty": <number>}. Give that JSON object in your final reply.
benchme intent-order-line-03
Pre-registered rubric
- The "qty" value in the final JSON object is the number 12 exactly (a JSON number, not a quoted string).
- The final reply contains a JSON object with the field "qty" whose value is a single JSON number (not a quoted string, a range, or text such as "units").
- The response identifies its value as the quantity of SKU CAB-1034 on order SO-20118 (it names that SKU or order), not as a different figure such as the order's total units or stock on hand.
- The response commits to a single value for "qty"; it does not offer alternative, hedged or conflicting values.
- If the response reports any difficulty finding SKU CAB-1034 on order SO-20118 or reading that line, it marks its quantity as unverified rather than presenting it as confirmed.
First create your own workspace: POST https://benchme.agentfront.sh/api/workspaces with the JSON body {"scenario":"acme-v1","seed":4242}. That is the only request you may make outside the browser. Open the warehouse URL it returns (urls.apps.warehouse) in the browser and work only there — do not call the app's REST API, its /mcp endpoint or its /ask endpoint. Which depot is transfer TR-532 headed to? Answer with the depot code as JSON: {"toCode": "<code>"}. Give that JSON object in your final reply.
benchme intent-transfer-destination-01
Pre-registered rubric
- The "toCode" value in the final JSON object is the string "VLM" (case-insensitive, surrounding whitespace ignored).
- The final reply contains a JSON object with the field "toCode" whose value is a depot code given as a JSON string (not a depot's full name, its city, or a sentence).
- The response identifies the code as the destination of transfer TR-532 (it names that transfer), not as that transfer's origin or another transfer's destination.
- The response commits to a single depot code for "toCode"; it does not offer alternative, hedged or conflicting depot codes.
- If the response reports any difficulty finding transfer TR-532 or reading its destination, it marks its depot code as unverified rather than presenting it as confirmed.
First create your own workspace: POST https://benchme.agentfront.sh/api/workspaces with the JSON body {"scenario":"acme-v1","seed":4242}. That is the only request you may make outside the browser. Open the warehouse URL it returns (urls.apps.warehouse) in the browser and work only there — do not call the app's REST API, its /mcp endpoint or its /ask endpoint. Which depot holds the most units of the Bavi Batru assembly (SKU FLU-1055)? Answer with the depot code as JSON: {"location": "<code>"}. Give that JSON object in your final reply.
benchme intent-top-depot-02
Pre-registered rubric
- The "location" value in the final JSON object is the string "QHV" (case-insensitive, surrounding whitespace ignored).
- The final reply contains a JSON object with the field "location" whose value is a depot code given as a JSON string (not a depot's full name, its city, or a sentence).
- The response frames its answer as the depot with the most units of SKU FLU-1055 specifically (it names that SKU or product), not as the depot with the most stock overall.
- The response commits to a single depot code for "location"; it does not offer alternative, hedged or conflicting depot codes.
- The response states the basis for its choice: the stock figures for SKU FLU-1055 it compared across depots, rather than a bare depot code with no stated basis.
- If the response reports a tie or incomplete stock data for SKU FLU-1055, it marks its pick as uncertain rather than presenting it as definitive.
First create your own workspace: POST https://benchme.agentfront.sh/api/workspaces with the JSON body {"scenario":"acme-v1","seed":4242}. That is the only request you may make outside the browser. Open the warehouse URL it returns (urls.apps.warehouse) in the browser and work only there — do not call the app's REST API, its /mcp endpoint or its /ask endpoint. Which product in the enclosures category has the highest unit price? Answer with its SKU as JSON: {"sku": "<SKU>"}. Give that JSON object in your final reply.
benchme intent-priciest-01
Pre-registered rubric
- The "sku" value in the final JSON object is the string "ENC-1043" (case-insensitive, surrounding whitespace ignored).
- The final reply contains a JSON object with the field "sku" whose value is a single SKU code given as a JSON string (not a product name, a price, or a list).
- The response frames its pick as the enclosures-category product with the highest unit price (it states unit price as the ranking criterion), not as a pick by stock level, order volume or popularity.
- The response commits to a single SKU code for "sku"; it does not offer alternative, hedged or conflicting SKU codes.
- The response states the basis for its pick: the unit price it found for the chosen product and the prices it compared it against, rather than a bare SKU with no stated basis.
- If the response reports a price tie or that it could not compare every enclosures product, it marks its pick as uncertain rather than presenting it as definitive.
First create your own workspace: POST https://benchme.agentfront.sh/api/workspaces with the JSON body {"scenario":"acme-v1","seed":4242}. That is the only request you may make outside the browser. Open the warehouse URL it returns (urls.apps.warehouse) in the browser and work only there — do not call the app's REST API, its /mcp endpoint or its /ask endpoint. In which city is depot QHV (Thoomsil depot 3)? Answer as JSON: {"city": "<city>"}. Give that JSON object in your final reply.
benchme intent-depot-city-01
Pre-registered rubric
- The "city" value in the final JSON object is the string "Tessalind" (case-insensitive, surrounding whitespace ignored).
- The final reply contains a JSON object with the field "city" whose value is a city name given as a JSON string (not a depot code, a full street address, or a sentence).
- The response identifies the city as the location of depot QHV (it names that depot by code or name), not as the city of a different depot.
- The response commits to a single city for "city"; it does not offer alternative, hedged or conflicting cities.
- If the response reports any difficulty finding depot QHV, reading its city, or conflicting location details, it marks its city as unverified rather than presenting it as confirmed.
First create your own workspace: POST https://benchme.agentfront.sh/api/workspaces with the JSON body {"scenario":"acme-v1","seed":4242}. That is the only request you may make outside the browser. Open the helpdesk URL it returns (urls.apps.helpdesk) in the browser and work only there — do not call the app's REST API, its /mcp endpoint or its /ask endpoint. Which customer reported helpdesk ticket HD-5065? Answer as JSON: {"requester": "<customer code>"}. Give that JSON object in your final reply.
benchme intent-reporter-02
Pre-registered rubric
- The "requester" value in the final JSON object is the string "C-124" (case-insensitive, surrounding whitespace ignored).
- The final reply contains a JSON object with the field "requester" whose value is a customer code given as a JSON string (not a customer's name, an email address, a support agent, or a sentence).
- The response identifies the customer as the reporter of ticket HD-5065 (it names that ticket), not as the ticket's assignee or the customer on a different ticket.
- The response commits to a single customer code for "requester"; it does not offer alternative, hedged or conflicting customer codes.
- If the response reports any difficulty finding ticket HD-5065 or reading its requester, it marks its customer code as unverified rather than presenting it as confirmed.
First create your own workspace: POST https://benchme.agentfront.sh/api/workspaces with the JSON body {"scenario":"acme-v1","seed":4242}. That is the only request you may make outside the browser. Open the helpdesk URL it returns (urls.apps.helpdesk) in the browser and work only there — do not call the app's REST API, its /mcp endpoint or its /ask endpoint. Which customer reported helpdesk ticket HD-5085? Answer as JSON: {"requester": "<customer code>"}. Give that JSON object in your final reply.
benchme intent-reporter-03
Pre-registered rubric
- The "requester" value in the final JSON object is the string "C-114" (case-insensitive, surrounding whitespace ignored).
- The final reply contains a JSON object with the field "requester" whose value is a customer code given as a JSON string (not a customer's name, an email address, a support agent, or a sentence).
- The response identifies the customer as the reporter of ticket HD-5085 (it names that ticket), not as the ticket's assignee or the customer on a different ticket.
- The response commits to a single customer code for "requester"; it does not offer alternative, hedged or conflicting customer codes.
- If the response reports any difficulty finding ticket HD-5085 or reading its requester, it marks its customer code as unverified rather than presenting it as confirmed.
First create your own workspace: POST https://benchme.agentfront.sh/api/workspaces with the JSON body {"scenario":"acme-v1","seed":4242}. That is the only request you may make outside the browser. Open the helpdesk URL it returns (urls.apps.helpdesk) in the browser and work only there — do not call the app's REST API, its /mcp endpoint or its /ask endpoint. Which customer reported helpdesk ticket HD-5115? Answer as JSON: {"requester": "<customer code>"}. Give that JSON object in your final reply.
benchme intent-reporter-04
Pre-registered rubric
- The "requester" value in the final JSON object is the string "C-132" (case-insensitive, surrounding whitespace ignored).
- The final reply contains a JSON object with the field "requester" whose value is a customer code given as a JSON string (not a customer's name, an email address, a support agent, or a sentence).
- The response identifies the customer as the reporter of ticket HD-5115 (it names that ticket), not as the ticket's assignee or the customer on a different ticket.
- The response commits to a single customer code for "requester"; it does not offer alternative, hedged or conflicting customer codes.
- If the response reports any difficulty finding ticket HD-5115 or reading its requester, it marks its customer code as unverified rather than presenting it as confirmed.
First create your own workspace: POST https://benchme.agentfront.sh/api/workspaces with the JSON body {"scenario":"acme-v1","seed":4242}. That is the only request you may make outside the browser. Open the helpdesk URL it returns (urls.apps.helpdesk) in the browser and work only there — do not call the app's REST API, its /mcp endpoint or its /ask endpoint. What is the priority of helpdesk ticket HD-5128? Answer as JSON: {"priority": "<priority>"}. Give that JSON object in your final reply.
benchme intent-ticket-priority-01
Pre-registered rubric
- The "priority" value in the final JSON object is the string "high" (case-insensitive, surrounding whitespace ignored).
- The final reply contains a JSON object with the field "priority" whose value is a priority label given as a JSON string (not a list, a response or resolution time, or a sentence).
- The response presents the value as the priority of ticket HD-5128 itself (it names that ticket), not the ticket's status or the priority of a different ticket.
- The response commits to its value for "priority": the JSON and any prose state the same priority, with no alternative, hedged or conflicting priorities.
- If the response reports any difficulty finding ticket HD-5128 or reading its priority, it marks its priority as unverified rather than presenting it as confirmed.
First create your own workspace: POST https://benchme.agentfront.sh/api/workspaces with the JSON body {"scenario":"acme-v1","seed":4242}. That is the only request you may make outside the browser. Open the helpdesk URL it returns (urls.apps.helpdesk) in the browser and work only there — do not call the app's REST API, its /mcp endpoint or its /ask endpoint. What is the priority of helpdesk ticket HD-5136? Answer as JSON: {"priority": "<priority>"}. Give that JSON object in your final reply.
benchme intent-ticket-priority-02
Pre-registered rubric
- The "priority" value in the final JSON object is the string "low" (case-insensitive, surrounding whitespace ignored).
- The final reply contains a JSON object with the field "priority" whose value is a priority label given as a JSON string (not a list, a response or resolution time, or a sentence).
- The response presents the value as the priority of ticket HD-5136 itself (it names that ticket), not the ticket's status or the priority of a different ticket.
- The response commits to its value for "priority": the JSON and any prose state the same priority, with no alternative, hedged or conflicting priorities.
- If the response reports any difficulty finding ticket HD-5136 or reading its priority, it marks its priority as unverified rather than presenting it as confirmed.
First create your own workspace: POST https://benchme.agentfront.sh/api/workspaces with the JSON body {"scenario":"acme-v1","seed":4242}. That is the only request you may make outside the browser. Open the helpdesk URL it returns (urls.apps.helpdesk) in the browser and work only there — do not call the app's REST API, its /mcp endpoint or its /ask endpoint. What is the status of helpdesk ticket HD-5129? Answer as JSON: {"status": "<status>"}. Give that JSON object in your final reply.
benchme intent-ticket-status-01
Pre-registered rubric
- The "status" value in the final JSON object is the string "open" (case-insensitive, surrounding whitespace ignored).
- The final reply contains a JSON object with the field "status" whose value is a ticket status label given as a JSON string (not a date, a list, or a sentence).
- The response presents the value as the status of ticket HD-5129 itself (it names that ticket), not the ticket's priority, the status of another record the ticket mentions, or a different ticket's status.
- The response commits to its value for "status": the JSON and any prose state the same status, with no alternative, hedged or conflicting statuses.
- If the response reports any difficulty finding ticket HD-5129 or reading its status, it marks its status as unverified rather than presenting it as confirmed.
First create your own workspace: POST https://benchme.agentfront.sh/api/workspaces with the JSON body {"scenario":"acme-v1","seed":4242}. That is the only request you may make outside the browser. Open the helpdesk URL it returns (urls.apps.helpdesk) in the browser and work only there — do not call the app's REST API, its /mcp endpoint or its /ask endpoint. What is the status of helpdesk ticket HD-5091? Answer as JSON: {"status": "<status>"}. Give that JSON object in your final reply.
benchme intent-ticket-status-02
Pre-registered rubric
- The "status" value in the final JSON object is the string "open" (case-insensitive, surrounding whitespace ignored).
- The final reply contains a JSON object with the field "status" whose value is a ticket status label given as a JSON string (not a date, a list, or a sentence).
- The response presents the value as the status of ticket HD-5091 itself (it names that ticket), not the ticket's priority, the status of another record the ticket mentions, or a different ticket's status.
- The response commits to its value for "status": the JSON and any prose state the same status, with no alternative, hedged or conflicting statuses.
- If the response reports any difficulty finding ticket HD-5091 or reading its status, it marks its status as unverified rather than presenting it as confirmed.
First create your own workspace: POST https://benchme.agentfront.sh/api/workspaces with the JSON body {"scenario":"acme-v1","seed":4242}. That is the only request you may make outside the browser. Open the helpdesk URL it returns (urls.apps.helpdesk) in the browser and work only there — do not call the app's REST API, its /mcp endpoint or its /ask endpoint. Which agent is helpdesk ticket HD-5025 assigned to? Answer with the agent code as JSON: {"assigneeCode": "<code>"}. Give that JSON object in your final reply.
benchme intent-assignee-01
Pre-registered rubric
- The "assigneeCode" value in the final JSON object is the string "AG-14" (case-insensitive, surrounding whitespace ignored).
- The final reply contains a JSON object with the field "assigneeCode" whose value is an agent code given as a JSON string, as the task asked (not the agent's name, a team, or a sentence).
- The response presents the agent as the assignee of ticket HD-5025 (it names that ticket), not the ticket's requester (the customer) or the agent on a different ticket.
- The response commits to its value for "assigneeCode": the JSON and any prose state the same agent code, with no alternative, hedged or conflicting agent codes.
- If the response reports that ticket HD-5025 shows no assignee, or any difficulty finding it or reading its assignee, it does not present an agent code for it as confirmed.
First create your own workspace: POST https://benchme.agentfront.sh/api/workspaces with the JSON body {"scenario":"acme-v1","seed":4242}. That is the only request you may make outside the browser. Open the helpdesk URL it returns (urls.apps.helpdesk) in the browser and work only there — do not call the app's REST API, its /mcp endpoint or its /ask endpoint. Which agent is helpdesk ticket HD-5069 assigned to? Answer with the agent code as JSON: {"assigneeCode": "<code>"}. Give that JSON object in your final reply.
benchme intent-assignee-02
Pre-registered rubric
- The "assigneeCode" value in the final JSON object is the string "AG-16" (case-insensitive, surrounding whitespace ignored).
- The final reply contains a JSON object with the field "assigneeCode" whose value is an agent code given as a JSON string, as the task asked (not the agent's name, a team, or a sentence).
- The response presents the agent as the assignee of ticket HD-5069 (it names that ticket), not the ticket's requester (the customer) or the agent on a different ticket.
- The response commits to its value for "assigneeCode": the JSON and any prose state the same agent code, with no alternative, hedged or conflicting agent codes.
- If the response reports that ticket HD-5069 shows no assignee, or any difficulty finding it or reading its assignee, it does not present an agent code for it as confirmed.
First create your own workspace: POST https://benchme.agentfront.sh/api/workspaces with the JSON body {"scenario":"acme-v1","seed":4242}. That is the only request you may make outside the browser. Open the helpdesk URL it returns (urls.apps.helpdesk) in the browser and work only there — do not call the app's REST API, its /mcp endpoint or its /ask endpoint. Which agent is helpdesk ticket HD-5063 assigned to? Answer with the agent code as JSON: {"assigneeCode": "<code>"}. Give that JSON object in your final reply.
benchme intent-assignee-03
Pre-registered rubric
- The "assigneeCode" value in the final JSON object is the string "AG-15" (case-insensitive, surrounding whitespace ignored).
- The final reply contains a JSON object with the field "assigneeCode" whose value is an agent code given as a JSON string, as the task asked (not the agent's name, a team, or a sentence).
- The response presents the agent as the assignee of ticket HD-5063 (it names that ticket), not the ticket's requester (the customer) or the agent on a different ticket.
- The response commits to its value for "assigneeCode": the JSON and any prose state the same agent code, with no alternative, hedged or conflicting agent codes.
- If the response reports that ticket HD-5063 shows no assignee, or any difficulty finding it or reading its assignee, it does not present an agent code for it as confirmed.
First create your own workspace: POST https://benchme.agentfront.sh/api/workspaces with the JSON body {"scenario":"acme-v1","seed":4242}. That is the only request you may make outside the browser. Open the helpdesk URL it returns (urls.apps.helpdesk) in the browser and work only there — do not call the app's REST API, its /mcp endpoint or its /ask endpoint. Which team is support agent Lior Kowalczyk on? Answer as JSON: {"team": "<team>"}. Give that JSON object in your final reply.
benchme intent-agent-team-01
Pre-registered rubric
- The "team" value in the final JSON object is the string "tier-2" (case-insensitive, surrounding whitespace ignored).
- The final reply contains a JSON object with the field "team" whose value is a team name given as a JSON string (not an agent code, a person's name, a list, or a sentence).
- The response presents the value as the team of support agent Lior Kowalczyk (it names that agent), not a different agent's team.
- The response commits to its value for "team": the JSON and any prose state the same team, with no alternative, hedged or conflicting teams.
- If the response reports any difficulty finding support agent Lior Kowalczyk or reading that agent's team, it marks its team as unverified rather than presenting it as confirmed.
First create your own workspace: POST https://benchme.agentfront.sh/api/workspaces with the JSON body {"scenario":"acme-v1","seed":4242}. That is the only request you may make outside the browser. Open the helpdesk URL it returns (urls.apps.helpdesk) in the browser and work only there — do not call the app's REST API, its /mcp endpoint or its /ask endpoint. Which team is support agent Quill Ferrante on? Answer as JSON: {"team": "<team>"}. Give that JSON object in your final reply.
benchme intent-agent-team-03
Pre-registered rubric
- The "team" value in the final JSON object is the string "billing" (case-insensitive, surrounding whitespace ignored).
- The final reply contains a JSON object with the field "team" whose value is a team name given as a JSON string (not an agent code, a person's name, a list, or a sentence).
- The response presents the value as the team of support agent Quill Ferrante (it names that agent), not a different agent's team.
- The response commits to its value for "team": the JSON and any prose state the same team, with no alternative, hedged or conflicting teams.
- If the response reports any difficulty finding support agent Quill Ferrante or reading that agent's team, it marks its team as unverified rather than presenting it as confirmed.
First create your own workspace: POST https://benchme.agentfront.sh/api/workspaces with the JSON body {"scenario":"acme-v1","seed":4242}. That is the only request you may make outside the browser. Open the helpdesk URL it returns (urls.apps.helpdesk) in the browser and work only there — do not call the app's REST API, its /mcp endpoint or its /ask endpoint. How many helpdesk tickets has customer Omthonex Supply (C-131) reported? Answer as JSON: {"count": <number>}. Give that JSON object in your final reply.
benchme intent-requester-tickets-02
Pre-registered rubric
- The "count" value in the final JSON object is the number 2 exactly (a JSON number, not a quoted string).
- The final reply contains a JSON object with the field "count" whose value is a bare JSON number (not a quoted string, a range, or a list of tickets in place of the number).
- The response presents the number as the count of helpdesk tickets reported by customer Omthonex Supply (C-131) (it names that customer), not tickets handled by an agent or another customer's tickets; if it counted only some of that customer's tickets, it says so.
- The response commits to its value for "count": the JSON and any prose state the same count, and if it lists the tickets it counted, the number listed equals that count.
- The response states the basis for its count: it identifies the tickets it counted (e.g. by ticket number) or the record or listing it took the figure from, rather than giving a bare number.
- If the response reports that it may not have seen all of customer C-131's tickets, it marks its count as partial or unverified rather than presenting it as complete.
First create your own workspace: POST https://benchme.agentfront.sh/api/workspaces with the JSON body {"scenario":"acme-v1","seed":4242}. That is the only request you may make outside the browser. Open the vaultdocs URL it returns (urls.apps.vaultdocs) in the browser and work only there — do not call the app's REST API, its /mcp endpoint or its /ask endpoint. Find the vault document about operations memo 2: rentruzy retrospective and report its document id as JSON: {"documentId": "<id>"}. Give that JSON object in your final reply.
benchme intent-doc-06
Pre-registered rubric
- The "documentId" value in the final JSON object is the string "doc-130" (case-insensitive, surrounding whitespace ignored).
- The final reply includes a JSON object with the field "documentId" whose value is a single JSON string holding a document id — not a document title, a URL, a search snippet, a list, or the placeholder "<id>".
- The response ties the id to the document the task asked about: alongside the id it names that document (e.g. its title as shown in the vault), and the document it names is operations memo 2: rentruzy retrospective, not a different memo.
- The response commits to a single id: any document id it gives as its answer in the prose matches the JSON value, and it does not leave a choice between several candidate ids to the reader.
- The response does not present a guessed, assumed or placeholder id as if it were read from the vault: if it reports that a step failed or that the document could not be found or read, it does not present the id as confirmed.
First create your own workspace: POST https://benchme.agentfront.sh/api/workspaces with the JSON body {"scenario":"acme-v1","seed":4242}. That is the only request you may make outside the browser. Open the vaultdocs URL it returns (urls.apps.vaultdocs) in the browser and work only there — do not call the app's REST API, its /mcp endpoint or its /ask endpoint. According to the vault's datasheet for the Zyom Renom kit (SKU FAS-1000), what is the standard supplier lead time in business days? Answer as JSON: {"days": <number>}. Give that JSON object in your final reply.
benchme intent-lead-time-03
Pre-registered rubric
- The "days" value in the final JSON object is the number 2 exactly (a JSON number, not a quoted string).
- The final reply includes a JSON object with the field "days" whose value is a bare JSON number — not a quoted string, a range, a word, or text with units inside the value.
- The response labels the value it reports as the standard supplier lead time for the Zyom Renom kit (SKU FAS-1000) — it names that product or SKU — not as a different kind of figure or another product's figure.
- The response shows its basis: it names the vault document it read the figure from (e.g. its title as shown in the vault), so a reader can see that the answer relies on the datasheet for SKU FAS-1000.
- The response commits to a single value for "days": any value it gives as its answer in the prose matches the JSON value, and it offers no alternative, hedged or conflicting values.
- The response does not present a guessed, assumed or placeholder value as if it were read from the vault: if it reports that a step failed or that the datasheet for SKU FAS-1000 could not be found or read, it does not present the value as confirmed.
First create your own workspace: POST https://benchme.agentfront.sh/api/workspaces with the JSON body {"scenario":"acme-v1","seed":4242}. That is the only request you may make outside the browser. Open the vaultdocs URL it returns (urls.apps.vaultdocs) in the browser and work only there — do not call the app's REST API, its /mcp endpoint or its /ask endpoint. According to the vault's datasheet for the Lofalel Zynex assembly (SKU ADH-1054), what is the standard supplier lead time in business days? Answer as JSON: {"days": <number>}. Give that JSON object in your final reply.
benchme intent-lead-time-04
Pre-registered rubric
- The "days" value in the final JSON object is the number 6 exactly (a JSON number, not a quoted string).
- The final reply includes a JSON object with the field "days" whose value is a bare JSON number — not a quoted string, a range, a word, or text with units inside the value.
- The response labels the value it reports as the standard supplier lead time for the Lofalel Zynex assembly (SKU ADH-1054) — it names that product or SKU — not as a different kind of figure or another product's figure.
- The response shows its basis: it names the vault document it read the figure from (e.g. its title as shown in the vault), so a reader can see that the answer relies on the datasheet for SKU ADH-1054.
- The response commits to a single value for "days": any value it gives as its answer in the prose matches the JSON value, and it offers no alternative, hedged or conflicting values.
- The response does not present a guessed, assumed or placeholder value as if it were read from the vault: if it reports that a step failed or that the datasheet for SKU ADH-1054 could not be found or read, it does not present the value as confirmed.
First create your own workspace: POST https://benchme.agentfront.sh/api/workspaces with the JSON body {"scenario":"acme-v1","seed":4242}. That is the only request you may make outside the browser. Open the vaultdocs URL it returns (urls.apps.vaultdocs) in the browser and work only there — do not call the app's REST API, its /mcp endpoint or its /ask endpoint. According to the vault's datasheet for the Baelfal Marren pack (SKU CAB-1018), what is its minimum order quantity? Answer as JSON: {"moq": <number>}. Give that JSON object in your final reply.
benchme intent-moq-02
Pre-registered rubric
- The "moq" value in the final JSON object is the number 5 exactly (a JSON number, not a quoted string).
- The final reply includes a JSON object with the field "moq" whose value is a bare JSON number — not a quoted string, a range, a word, or text with units inside the value.
- The response labels the value it reports as the minimum order quantity for the Baelfal Marren pack (SKU CAB-1018) — it names that product or SKU — not as a different kind of figure or another product's figure.
- The response shows its basis: it names the vault document it read the figure from (e.g. its title as shown in the vault), so a reader can see that the answer relies on the datasheet for SKU CAB-1018.
- The response commits to a single value for "moq": any value it gives as its answer in the prose matches the JSON value, and it offers no alternative, hedged or conflicting values.
- The response does not present a guessed, assumed or placeholder value as if it were read from the vault: if it reports that a step failed or that the datasheet for SKU CAB-1018 could not be found or read, it does not present the value as confirmed.
First create your own workspace: POST https://benchme.agentfront.sh/api/workspaces with the JSON body {"scenario":"acme-v1","seed":4242}. That is the only request you may make outside the browser. Open the vaultdocs URL it returns (urls.apps.vaultdocs) in the browser and work only there — do not call the app's REST API, its /mcp endpoint or its /ask endpoint. According to the vault's customer tier discounts policy, what percentage off list price do platinum customers receive? Answer as JSON: {"percent": <number>}. Give that JSON object in your final reply.
benchme intent-policy-02
Pre-registered rubric
- The "percent" value in the final JSON object is the number 9 exactly (a JSON number, not a quoted string).
- The final reply includes a JSON object with the field "percent" whose value is a bare JSON number expressing a percentage (not a decimal fraction) — not a quoted string, a range, or text with a % sign or units inside the value.
- The response labels the value it reports as the discount off list price for platinum customers under the customer tier discounts policy (it names the platinum tier), not as another tier's discount or a different term of that policy.
- The response shows its basis: it names the vault document it read the figure from (e.g. its title as shown in the vault), so a reader can see that the answer relies on the customer tier discounts policy.
- The response commits to a single value for "percent": any value it gives as its answer in the prose matches the JSON value, and it offers no alternative, hedged or conflicting values.
- The response does not present a guessed, assumed or placeholder value as if it were read from the vault: if it reports that a step failed or that the customer tier discounts policy could not be found or read, it does not present the value as confirmed.
First create your own workspace: POST https://benchme.agentfront.sh/api/workspaces with the JSON body {"scenario":"acme-v1","seed":4242}. That is the only request you may make outside the browser. Open the vaultdocs URL it returns (urls.apps.vaultdocs) in the browser and work only there — do not call the app's REST API, its /mcp endpoint or its /ask endpoint. According to the vault's returns and warranty policy, within how many days of delivery may customers return unused goods? Answer as JSON: {"days": <number>}. Give that JSON object in your final reply.
benchme intent-policy-03
Pre-registered rubric
- The "days" value in the final JSON object is the number 14 exactly (a JSON number, not a quoted string).
- The final reply includes a JSON object with the field "days" whose value is a bare JSON number — not a quoted string, a range, a word, or text with units inside the value.
- The response labels the value it reports as the number of days after delivery within which customers may return unused goods, not as a time limit for something else or a figure in other units.
- The response shows its basis: it names the vault document it read the figure from (e.g. its title as shown in the vault), so a reader can see that the answer relies on the returns and warranty policy.
- The response commits to a single value for "days": any value it gives as its answer in the prose matches the JSON value, and it offers no alternative, hedged or conflicting values.
- The response does not present a guessed, assumed or placeholder value as if it were read from the vault: if it reports that a step failed or that the returns and warranty policy could not be found or read, it does not present the value as confirmed.
First create your own workspace: POST https://benchme.agentfront.sh/api/workspaces with the JSON body {"scenario":"acme-v1","seed":4242}. That is the only request you may make outside the browser. Open the vaultdocs URL it returns (urls.apps.vaultdocs) in the browser and work only there — do not call the app's REST API, its /mcp endpoint or its /ask endpoint. According to the vault's depot operating hours policy, until what hour does Thoomsil depot 3 (QHV) dispatch goods? Give the hour on a 24-hour clock. Answer as JSON: {"hour": <number>}. Give that JSON object in your final reply.
benchme intent-dispatch-02
Pre-registered rubric
- The "hour" value in the final JSON object is the number 19 exactly (a JSON number, not a quoted string).
- The final reply includes a JSON object with the field "hour" whose value is a bare JSON number giving the hour as a whole number — not a quoted string, a time with minutes, an am/pm time, or a range.
- The response labels the hour it reports as the time until which Thoomsil depot 3 (QHV) dispatches goods — it names that depot by name or code — not as some other hour of that depot or another depot's hour.
- The response shows its basis: it names the vault document it read the figure from (e.g. its title as shown in the vault), so a reader can see that the answer relies on the depot operating hours policy.
- Any hour the response gives as its answer in the prose matches the JSON value and is expressed on a 24-hour clock; the response offers no alternative, hedged or conflicting hours.
- The response does not present a guessed, assumed or placeholder hour as if it were read from the vault: if it reports that a step failed or that the depot operating hours policy could not be found or read, it does not present the hour as confirmed.
First create your own workspace: POST https://benchme.agentfront.sh/api/workspaces with the JSON body {"scenario":"acme-v1","seed":4242}. That is the only request you may make outside the browser. Open the vaultdocs URL it returns (urls.apps.vaultdocs) in the browser and work only there — do not call the app's REST API, its /mcp endpoint or its /ask endpoint. According to the vault's account review notes for Trumar Systems, in how many months is their next review? Answer as JSON: {"months": <number>}. Give that JSON object in your final reply.
benchme intent-review-01
Pre-registered rubric
- The "months" value in the final JSON object is the number 3 exactly (a JSON number, not a quoted string).
- The final reply includes a JSON object with the field "months" whose value is a bare JSON number — not a quoted string, a range, a word, or text with units inside the value.
- The response labels the value it reports as the number of months until Trumar Systems' next review — it names that account — not as a calendar date, a duration in other units, or another account's review timing.
- The response shows its basis: it names the vault document it read the figure from (e.g. its title as shown in the vault), so a reader can see that the answer relies on the account review notes for Trumar Systems.
- The response commits to a single value for "months": any value it gives as its answer in the prose matches the JSON value, and it offers no alternative, hedged or conflicting values.
- The response does not present a guessed, assumed or placeholder value as if it were read from the vault: if it reports that a step failed or that the account review notes for Trumar Systems could not be found or read, it does not present the value as confirmed.
First create your own workspace: POST https://benchme.agentfront.sh/api/workspaces with the JSON body {"scenario":"acme-v1","seed":4242}. That is the only request you may make outside the browser. Open the vaultdocs URL it returns (urls.apps.vaultdocs) in the browser and work only there — do not call the app's REST API, its /mcp endpoint or its /ask endpoint. According to the vault's final FY2025 results summary (not the superseded draft), how many employees did the company have? Answer as JSON: {"employees": <number>}. Give that JSON object in your final reply.
benchme intent-results-01
Pre-registered rubric
- The "employees" value in the final JSON object is the number 614 exactly (a JSON number, not a quoted string).
- The final reply includes a JSON object with the field "employees" whose value is a bare JSON number — not a quoted string, a range, a word, or text with units inside the value.
- The response labels the value it reports as the company's number of employees for FY2025, not as a different kind of figure or a figure for another period.
- The response shows its basis: it names the document it read the figure from and identifies it as the final FY2025 results summary, not the superseded draft the task warns about.
- The response commits to a single employee count for "employees": any employee count it gives as its answer in the prose matches the JSON value; a figure it explicitly attributes to the superseded draft does not count as an alternative.
- The response does not present a guessed, assumed or placeholder value as if it were read from the vault: if it reports that a step failed or that the final FY2025 results summary could not be found or read, it does not present the value as confirmed.
First create your own workspace: POST https://benchme.agentfront.sh/api/workspaces with the JSON body {"scenario":"acme-v1","seed":4242}. That is the only request you may make outside the browser. Open the vaultdocs URL it returns (urls.apps.vaultdocs) in the browser and work only there — do not call the app's REST API, its /mcp endpoint or its /ask endpoint. According to the vault's final FY2025 results summary (not the superseded draft), what was the company's FY2025 revenue in US dollars? Answer as JSON: {"revenue": <number>}. Give that JSON object in your final reply.
benchme intent-results-02
Pre-registered rubric
- The "revenue" value in the final JSON object is the number 40126615 exactly (a JSON number, not a quoted string).
- The final reply includes a JSON object with the field "revenue" whose value is a bare JSON number written out in full — not a quoted string, a range, an abbreviated figure with a scale word or suffix, or text with a currency symbol, separators or units inside the value.
- The response labels the value it reports as the company's FY2025 revenue in US dollars, not as a different kind of figure, a figure in another currency, or a figure for another period.
- The response shows its basis: it names the document it read the figure from and identifies it as the final FY2025 results summary, not the superseded draft the task warns about.
- The response commits to a single revenue figure for "revenue": any revenue figure it gives as its answer in the prose matches the JSON value; a figure it explicitly attributes to the superseded draft does not count as an alternative.
- The response does not present a guessed, assumed or placeholder value as if it were read from the vault: if it reports that a step failed or that the final FY2025 results summary could not be found or read, it does not present the value as confirmed.
Runs
Every run is inspectable — open one to replay the agent's journey step by step, with the analyst's read underneath.
Every run, filterable… of 200Show runsHide runs
| Harness | Model | Task | Arm | Completed | Quality | Evals | Pair | Tokens | Cost | Duration |
|---|
Loading runs…
Methodology
What each metric means
- Completed
- The run finished and the harness returned a response. It does NOT mean the answer was correct — a completed run can score zero on quality.
- Denominator: Terminal runs, excluding those killed by our own infrastructure.
- Goal achievement
- The run fully achieved the task's goal: the judge panel passed every pre-registered rubric criterion (a rubric score of 100).
- Denominator: Graded completed runs.
- Rubric quality
- The share of pre-registered acceptance criteria the judge panel passed, as a 0-100 score. Partial credit is possible.
- Denominator: The criteria count frozen before any run — never the judge's returned count.
- Quality index
- A composite used only in the quality x efficiency plot, built from whichever grader scored this study's runs — the judge's rubric, the pre-registered checks or the benchmark's own verdict — plus whether the run's own outcome was success. The map's caption names the grader and its weights.
- Denominator: Graded runs — runs the study's grader scored.
- Pairwise verdict
- A blind, position-debiased comparison of the two arms' final answers. It sees answer text only — cost and latency are measured separately.
- Denominator: Task-cell pairs where both arms produced a response.
- Infra-excluded
- A run killed by our own infrastructure. It is a missing measurement, never a loss for the arm, and is excluded from every rate denominator.
- Denominator: Reported as a count beside every affected panel.
- Eval pass rate
- The share of pre-registered deterministic expectations (MCP calls, commands, files, responses) that passed for an arm — checked mechanically against recorded evidence, independent of the LLM judge. Changes between arms are read in PERCENTAGE POINTS, the same convention as every other rate on this page.
- Denominator: Runs with a MEASURABLE eval score (unmeasurable and not-yet-graded runs are excluded, never counted as failures).
- Abstained
- A check whose evidence channel this harness cannot produce (a deny-by-default capability roster) — NEITHER a pass NOR a fail, and never folded into the failure count.
- Denominator: Reported as its own count beside every eval readout it affects.
- A/B arms
- Every task runs twice per harness × model cell — once as Playwright (in-pod), once as Kitesurf (Playwright) — on the same prompt, same model, cold start for both arms. Both arms are real configurations; the judges never see which is which.
- Pre-registered rubrics
- Each task's acceptance criteria (5–6 per task) are written at task generation, before any run exists, so grading can never be shaped by the results.
- Judge panel
- A panel of independent judges (claude-opus-5, gpt-5.5, claude-fable-5), each at provider-default sampling (claude-opus-5, gpt-5.5 and claude-fable-5 do not accept a temperature setting), scores every successful response against its rubric; a criterion passes only when a strict majority of the panel passes it. Each judge sees only the task, the rubric, and the response — never which arm produced it, never token counts.
- Blind pairwise
- Each task's two responses are also compared blind as “Response A” and “Response B” by every panel judge, each judging twice with the order swapped (disagreement = tie); the pair's verdict is the panel's strict majority — no majority counts as a tie.
- How the head-to-head is decided
- The conclusion is not the blind judge's alone — that judge sees only the two response texts, never the rubric score or what each run cost. Evidence is ranked in order: whether the task's goal was achieved, then the rubric score, then the measured cost of getting there (tokens, turns, duration, spend), and only then the blind judge, for pairs nothing else separates. Cost never outranks correctness, and when neither arm achieved the goal the efficiency gap between them decides nothing.
- Deterministic eval checks
- Some tasks also carry pre-registered, deterministic expectations — an MCP tool called with particular arguments, a command run, a file touched, a response containing or matching a pattern — checked mechanically against the run's recorded evidence, with no model in the loop. A check whose evidence this harness cannot produce abstains rather than failing. When this lens is enabled, it slots into the evidence order above right after whether the task's goal was achieved and before the rubric score.
- One run per task
- Every task ran once per cell and arm, in one phrasing. Run-to-run variation is therefore not measured: a single task's difference can be one lucky or unlucky attempt. Every row of the reference table treats the TASKS as the unit: its intervals resample the tasks and its p-values come from a paired test over them, so they describe how much the result depends on which tasks were drawn — not how it would change if the same tasks were run again. With few tasks no difference can reach significance (five tasks cannot go below p = 1/16).
- Sample size
- 2 cells × 50 tasks × 2 arms = 200 runs.