96%
Puzzle Box is the outlier.
GPT-6 Astra reaches 96% success, compared with 70% for Claude Fable 5.1 and 12% for Gemini 3.8 Flash. Excluding this task lowers their means to 10.9%, 3.3% and 0.9%.
Results
All three frontier agents play the same 50 instances per task. Even the strongest reaches 19.4% success overall, despite 44.9% mean task progress. Progress is not completion.
Equal-weight mean · 500 episodes per model
Equal-weight mean · 450 episodes per model
| # | Model | Reasoning | Success rate (%) | Progress (%) | Success w/o Puzzle Box (%) | $ / episode | $ / success |
|---|---|---|---|---|---|---|---|
| 1 | GPT-6 Astra | Medium thinking |
19.4
|
44.9 | 10.9 | 11.30 | 58.2 |
| 2 | Claude Fable 5.1 | Medium thinking |
10.0
|
29.3 | 3.3 | 19.22 | 192.2 |
| 3 | Gemini 3.8 Flash | High thinking |
2.0
|
9.2 | 0.9 | 3.37 | 168.3 |
| – | π₀.₅ (fine-tuned) | Trained VLA baseline, reported separately: nine tasks, a different protocol, and 0% success on eight of them. Baseline results | |||||
Per-task results
Switch between success rate and mean task progress, or leave out Puzzle Box, the one task on which all three models score far above their other tasks.
96%
GPT-6 Astra reaches 96% success, compared with 70% for Claude Fable 5.1 and 12% for Gemini 3.8 Flash. Excluding this task lowers their means to 10.9%, 3.3% and 0.9%.
0/150
None of the three models succeeds in its 50 evaluation episodes. Finding and positioning a light is only the beginning of the investigation.
Evaluation protocol
Frontier agents act through low-level arm and base tools, with no semantic object-location or grasp tools. A trained vision-language-action policy is reported separately.
Frontier agents
GPT-6 Astra and Claude Fable 5.1 use medium thinking; Gemini 3.8 Flash uses high thinking. Each decision receives three 512 × 512 images and proprioception.
Up to 200 decisions; up to 10 simulated seconds per command. The full text history is retained, with images from only the two latest observations.
Evaluation harness ↗Trained VLA baseline
RoboQuest tests generalist physical agents: a new task should not require a new fine-tuning run. π₀.₅ is evaluated with task-specific demonstration fine-tuning, yet the reported policy still falls short of the investigation and completion capabilities these tasks require.
These results expose a capability gap in the evaluated π₀.₅ setup, even with task-specific training. They do not establish that all VLA architectures are incapable of these tasks.
Training and evaluation details ↗Task-specific fine-tuning
π₀.₅ achieves 0% success on eight of nine reported tasks. Puzzle Box reaches 1.9% success despite 42.7% progress, exposing a substantial gap between partial action and task completion.
Nine reported tasks after task-specific fine-tuning. Progress reaches 3.8% on Search Room, 1.5% on Marked Mugs and 1.1% on Wobbly Stand; the remaining five tasks have 0% progress. Painted Cubes is not reported in the supplied figure. Per-task episode counts and uncertainty intervals are not provided; these results are separate from the frontier-agent evaluation.
| Task | Success (%) | Progress (%) |
|---|---|---|
| Puzzle Box | 1.9 | 42.7 |
| Search Room | 0.0 | 3.8 |
| Marked Mugs | 0.0 | 1.5 |
| Wobbly Stand | 0.0 | 1.1 |
| Locked Storage | 0.0 | 0.0 |
| Blackout Search | 0.0 | 0.0 |
| Unfamiliar Containers | 0.0 | 0.0 |
| Stamp Composition | 0.0 | 0.0 |
| Odd Parcel | 0.0 | 0.0 |
Cost and effort
Decisions, simulated and wall-clock time, tokens and list-price cost per episode for each frontier agent.
Costs are the manuscript’s recorded USD list-price accounting, not a current pricing quotation. Cost per success divides total spend by successful episodes.