RoboQuestby DeCLaRe Lab GitHub

Results

Leaderboard

All three frontier agents play the same 50 instances per task. Even the strongest reaches 19.4% success overall, despite 44.9% mean task progress. Progress is not completion.

All 10 tasks

Equal-weight mean · 500 episodes per model

Without Puzzle Box

Equal-weight mean · 450 episodes per model

Success and progress are equal-weight means over 10 tasks, 50 episodes per task and 500 per model; bars use a 0–100% scale. Costs are the manuscript’s recorded USD list-price accounting, not a current pricing quotation. Cost per success divides total spend by successful episodes.
# Model Reasoning Success rate (%) Progress (%) Success w/o Puzzle Box (%) $ / episode $ / success
1 GPT-6 Astra Medium thinking
19.4
44.9 10.9 11.30 58.2
2 Claude Fable 5.1 Medium thinking
10.0
29.3 3.3 19.22 192.2
3 Gemini 3.8 Flash High thinking
2.0
9.2 0.9 3.37 168.3
– π₀.₅ (fine-tuned) Trained VLA baseline, reported separately: nine tasks, a different protocol, and 0% success on eight of them. Baseline results

Per-task results

Results by task

Switch between success rate and mean task progress, or leave out Puzzle Box, the one task on which all three models score far above their other tasks.

Success rate by task

The interactive chart needs JavaScript. The full results table follows.

Equal-weight mean over 10 tasks · 500 episodes per model.Percent · 0–100

Success and task progress

Open table ↗
Main results. SR: success rate. Prog: mean progress. Both in %, 50 episodes per cell. Best SR per row in bold. Overall w/o Puzzle Box excludes the easiest task, an outlier on which all three models score far above their other tasks.
GPT-6 Astra Claude Fable 5.1 Gemini 3.8 Flash
Family Task SR Prog SR Prog SR Prog
Search Locked Storage 28.0 52.9 8.0 21.7 0.0 9.2
Search Room 0.0 31.0 0.0 19.2 2.0 17.5
Blackout Search 0.0 23.7 0.0 17.1 0.0 4.4
Inspect Painted Cubes 30.0 74.8 14.0 61.5 2.0 11.3
Marked Mugs 10.0 43.0 0.0 16.6 0.0 0.0
Unfamiliar Containers 6.0 32.2 0.0 17.0 0.0 6.3
Test Puzzle Box 96.0 99.6 70.0 81.0 12.0 33.0
Stamp Composition 14.0 55.3 2.0 36.9 0.0 3.4
Wobbly Stand 10.0 18.6 6.0 14.4 2.0 4.6
Odd Parcel 0.0 18.3 0.0 7.2 2.0 2.0
Overall 19.4 44.9 10.0 29.3 2.0 9.2
Overall w/o Puzzle Box 10.9 38.9 3.3 23.6 0.9 6.5

96%

Puzzle Box is the outlier.

GPT-6 Astra reaches 96% success, compared with 70% for Claude Fable 5.1 and 12% for Gemini 3.8 Flash. Excluding this task lowers their means to 10.9%, 3.3% and 0.9%.

0/150

Blackout Search remains unsolved.

None of the three models succeeds in its 50 evaluation episodes. Finding and positioning a light is only the beginning of the investigation.

Evaluation protocol

One decision. One physical command.

Frontier agents act through low-level arm and base tools, with no semantic object-location or grasp tools. A trained vision-language-action policy is reported separately.

Frontier agents

Same instances, same interface.

GPT-6 Astra and Claude Fable 5.1 use medium thinking; Gemini 3.8 Flash uses high thinking. Each decision receives three 512 × 512 images and proprioception.

Up to 200 decisions; up to 10 simulated seconds per command. The full text history is retained, with images from only the two latest observations.

Evaluation harness ↗

Trained VLA baseline

Task-specific training falls short.

RoboQuest tests generalist physical agents: a new task should not require a new fine-tuning run. π₀.₅ is evaluated with task-specific demonstration fine-tuning, yet the reported policy still falls short of the investigation and completion capabilities these tasks require.

These results expose a capability gap in the evaluated π₀.₅ setup, even with task-specific training. They do not establish that all VLA architectures are incapable of these tasks.

Training and evaluation details ↗

Task-specific fine-tuning

Training on the task is still not enough.

π₀.₅ achieves 0% success on eight of nine reported tasks. Puzzle Box reaches 1.9% success despite 42.7% progress, exposing a substantial gap between partial action and task completion.

π₀.₅ after task-specific fine-tuning

The chart needs JavaScript. Exact values are in the table below.

Separate from the frontier-agent evaluation.Percent · 0–50

Nine reported tasks after task-specific fine-tuning. Progress reaches 3.8% on Search Room, 1.5% on Marked Mugs and 1.1% on Wobbly Stand; the remaining five tasks have 0% progress. Painted Cubes is not reported in the supplied figure. Per-task episode counts and uncertainty intervals are not provided; these results are separate from the frontier-agent evaluation.

Exact fine-tuning results +
Fine-tuned π₀.₅ results from the updated author-supplied figure.
Task Success (%) Progress (%)
Puzzle Box 1.9 42.7
Search Room 0.0 3.8
Marked Mugs 0.0 1.5
Wobbly Stand 0.0 1.1
Locked Storage 0.0 0.0
Blackout Search 0.0 0.0
Unfamiliar Containers 0.0 0.0
Stamp Composition 0.0 0.0
Odd Parcel 0.0 0.0

Cost and effort

What an episode costs.

Decisions, simulated and wall-clock time, tokens and list-price cost per episode for each frontier agent.

Cost and interaction effort

Open table ↗
Cost and effort per episode. Sim / Wall: simulated / wall-clock minutes. In / Out: tokens in millions. Costs in USD at list prices.
Effort per episode Tokens per episode Cost (USD)
Model Decisions Sim min. Wall min. In (M) Out (M) Cached (%) $/episode $/success
GPT-6 Astra 143.7 10.4 53.8 5.82 0.02 91.8 11.30 58.2
Claude Fable 5.1 148.8 6.1 54.7 14.55 0.11 94.2 19.22 192.2
Gemini 3.8 Flash 160.8 5.1 41.0 19.66 0.13 89.3 3.37 168.3

Costs are the manuscript’s recorded USD list-price accounting, not a current pricing quotation. Cost per success divides total spend by successful episodes.