Search
Choose where to look. A negative finding redirects the search; a discovered token opens new possibilities.
Location uncertaintyRobots must acquire task-relevant evidence through physical interaction,
then decide when they know enough to commit.
DeCLaRe Lab · Goal-directed embodied exploration
Open storage places, track what has been checked, and retrieve the complete target set.
Turn objects to reveal unseen faces before deciding which bin they belong in.
Make trial impressions to discover patterns and rotations before committing permanent marks.
Selected demonstrations with all three camera views. Left scene camera · right scene camera · wrist camera. Aggregate results are reported separately.
01 / THE BENCHMARK
RoboQuest places a mobile manipulator in unfamiliar kitchens. The goal is given; task-critical information must be physically acquired.
Choose where to look. A negative finding redirects the search; a discovered token opens new possibilities.
Location uncertaintyLift, turn or open an object. Decide when the evidence is sufficient to identify a hidden property.
Hidden object propertiesIntervene and observe the outcome. Infer weight, a stamp pattern, a release order or an unseen tilt.
Unknown physical behaviorA physical button press commits the result and freezes the score. Every goal condition must hold. Stopping without pressing it is a failure.
A set of target items is stored in open or closed locations around a room, among distractor items of similar kind. The goal is to collect the complete set on the tray and press SUBMIT. The locations, the number of storage places, and which storage places are empty are not given. The robot must choose where to search, use negative findings to redirect, remember what it has checked, and stop once the set is complete. Instances vary the set size, item category, room layout, number and type of storage units, and distractors.
Storage compartments open only when a matching access token rests on their reader. Tokens are either outside the initial observation or hidden inside other storage, and some compartments contain tokens for other compartments, forming a dependency chain. The goal is to retrieve the target item. Finding a token changes which places can be searched next. Closing a compartment with a token still inside locks that token away permanently, so the robot must plan token handling as well as search order. Instances vary the number of locked compartments, chain structure, token hiding places, and dead ends.
A search task is performed in an unlit room. Only a light source with a glow marker and the SUBMIT button are visible at the start. The robot must acquire the light source, carry it, and search using a wrist camera that sees only what is illuminated. Because one gripper holds the light, retrieving a target requires putting the light down while aiming it at the target, then acting from memory in the dark. Instances vary the target, room, location of the light source, whether a wall switch exists, and whether a timed light turns off after a delay.
Several identical vessels, all mugs or all bowls in a given instance, each hold loose contents, and a label on the underside of each vessel specifies its destination. The goal is to place every vessel upright at its destination with its original contents inside. Reading a label requires lifting or tilting the vessel, which spills the contents; spilled contents stay nearby and must be returned to the correct vessel before submission. Instances vary the number of vessels, contents, label-to-destination mapping, and initial poses.
A collection of similar objects each carries a hidden property distributed over its faces, such as the number of painted faces, with some faces turned toward the table or away from every camera. The goal is to collect every object satisfying a stated property, for example, exactly one painted face. Partial views can reject an object but cannot confirm it, so the robot must decide how much inspection each object needs. Instances vary the number of objects, property rule, initial orientations, and whether the goal asks for one qualifying object or all of them.
Several closed containers each open by a different mechanism, such as a hinged lid, sliding lid, drawer, or latched flip-top, and some cannot be opened at all. The target is inside one of them. The robot knows neither where the target is nor how each container opens and must discover each mechanism by interaction. Instances vary the number of containers, mix of mechanisms, target location, and number of decoys.
Several stamps with unmarked housings each carry a hidden pattern and start at an unknown rotation. A target pattern is shown on a reference. A test surface allows trial impressions, and a final surface must end with exactly the target. The robot must test stamps, read the results, select a subset and their rotations, compose the target, and press SUBMIT. Marks on the final surface are permanent. Instances vary the number of stamps, patterns, target, initial rotations, and size of the test surface.
A set of visually identical sealed parcels contains one or two parcels that differ in weight. The only instrument is a two-pan balance whose pans hold several parcels each. The goal is to place the odd parcel alone on the tray. Group weighing identify it in few comparisons, while one-by-one weighing is legal but slower. Instances vary the number of parcels, whether the odd parcel is known to be heavier or could be lighter, pan capacity, and parcel appearance.
A container is held closed by a chain of interlocking parts, where each part blocks another until it is moved. All parts are visible with large handles, but the blocking relations are not. The goal is to place the item inside on the tray. The robot must discover the release order by trying moves and observing what moves, and some wrong moves jam another part until reversed. Instances vary the number of parts, chain structure, presence of red-herring parts, and trap moves.
A stand or table has an uneven support, tilting its surface by an amount too small to see directly. A ball placed on the surface rolls toward the low side and reveals the tilt. Shim blocks of different thicknesses are available. The goal is a level surface on which the ball stays put, with everything released. Instances vary the number of legs affected, tilt magnitude, shim set, and whether the tilt is along one axis or two.

A SHARED PHYSICAL WORLD
THE CAPABILITIES BEHIND THE TASKS
Ten tasks probe nine capabilities across four groups. The matrix marks direct targets and supporting requirements; success alone does not isolate any one capability.
Directed search chooses what to investigate.
Active perception makes hidden evidence visible.
Evidence integration combines observations.
Memory retains what is no longer in view.
Affordance discovery finds how objects work.
Experimental identification distinguishes hypotheses.
Physical causal inference relates actions to outcomes.
Consequential action accounts for irreversible effects.
Long-horizon planning orders dependent steps.
✓ Directly targeted · ∼ Supporting requirement · — No designated requirement. Capability links describe the task design, not independent causal measurements.
Feature designations reproduce the manuscript comparison; they describe designated benchmark targets, not the limits of every task in each suite.
02 / EVALUATION
All three frontier agents play the same 50 instances per task. Even the strongest reaches 19.4% success overall, despite 44.9% mean task progress.
Equal-weight mean over 10 tasks · 500 episodes per model · percentages.
GPT-6 Astra reaches 96% success, compared with 70% for Claude Fable 5.1 and 12% for Gemini 3.8 Flash. Excluding this task lowers their means to 10.9%, 3.3% and 0.9%.
None of the three models succeeds in its 50 evaluation episodes. Finding and positioning a light is only the beginning of the investigation.
FRONTIER AGENTS
GPT-6 Astra and Claude Fable 5.1 use medium thinking; Gemini 3.8 Flash uses high thinking. Each decision receives three 512 × 512 images and proprioception.
Up to 200 decisions; up to 10 simulated seconds per command. The full text history is retained, with images from only the two latest observations.
TRAINED VLA BASELINE
RoboQuest tests generalist physical agents: a new task should not require a new fine-tuning run. π₀.₅ is evaluated with task-specific demonstration fine-tuning, yet the reported policy still falls short of the investigation and completion capabilities these tasks require.
The reported five-task evaluation reaches 1.9% success and 42.7% progress on Puzzle Box. Marked Mugs, Wobbly Stand, Stamp Composition and Odd Parcel have 0% success.
These results expose a capability gap in the evaluated π₀.₅ setup, even with task-specific training. They do not establish that all VLA architectures are incapable of these tasks.
Training and evaluation details ↗Costs are the manuscript’s recorded USD list-price accounting, not a current pricing quotation. Cost per success divides total spend by successful episodes.
03 / FAILURE ANALYSIS
A failed episode can leave several requirements unmet. The analysis attributes each unmet requirement to the first uncorrected failure in the interaction sequence.
FAILURE TAXONOMY
Failures are organized by where the connection between evidence, decision and physical action first breaks and remains uncorrected.
Was the necessary evidence gathered, understood and used?
Did the command achieve the intended physical result?
Was the outcome checked, repaired if needed, and completion judged correctly?
| Phase | Failure category | What breaks | Example | Linked capabilities |
|---|---|---|---|---|
| Before action | Inadequate evidence gathering | Needed evidence was never acquired. | A location remains unchecked, a hidden face unobserved, or a diagnostic test omitted. | Evidence acquisition · Interactive inference |
| Before action | Misinterpretation of evidence | The evidence was observed but read incorrectly. | A label is misread or a trial outcome assigned the wrong meaning. | Evidence use |
| Before action | Faulty reasoning or planning | Available evidence leads to an invalid decision or plan. | Balance comparisons are combined incorrectly, or a dependency is ignored. | Evidence use · Interactive inference · Action organization |
| During action | Failed action execution | The intended physical outcome is not achieved. | A grasp misses, a carried object falls, or a shim lands beside its leg. | Action organization |
| After feedback | Failure to detect errors | A consequential error is not recognized. | A dropped object is left behind as the agent continues. | Evidence use · Action organization |
| After feedback | Ineffective error recovery | The agent attempts a repair but does not correct the error. | Repeated retrieval attempts leave the fallen object on the floor. | Interactive inference · Action organization |
| After feedback | Incorrect termination | The agent submits or stops with reachable goals unfinished. | A vessel is never handled before the robot ends the episode. | Evidence use · Action organization |
INADEQUATE EVIDENCE GATHERING
Share of unmet requirements in failed episodes. Denominators: 901 for Astra and 1,129 for Fable; 95% bootstrap intervals: [34, 40] and [39, 45].
In Search Room and Blackout Search, the gripper never approached within 25 cm of 83% of Astra’s missed targets and 91% of Fable’s. Relevant places remain unchecked.
From low to high workload, success falls by 15.8 percentage points for Astra and 7.2 for Fable on the four tasks that vary both workload and occlusion.
On the covered Puzzle Box subset, Astra reaches 96.4% success. Fable drops from 100% without the cover to 34.6% with it.
Failed execution accounts for 21.1% of Astra’s unmet requirements and 15.4% of Fable’s. Incorrect termination contributes 11.3% and 20.5%.
A separate analysis cohort. 540 matched instances per model, with 46–63 per task. This differs from the 500-episode main evaluation. The breakdown contains 432 failed episodes for Astra and 486 for Fable; six ambiguous Odd Parcel outcomes are excluded. These percentages are descriptive attributions, not isolated causal effects.
TASK-LEVEL DIAGNOSIS
Search failures are dominated by missing evidence. Marked Mugs and Wobbly Stand more often break during execution. Inspect each task’s distribution below.
Shares of unmet requirements within failed episodes in the matched analysis cohort. Displayed percentages are the reported rounded values; segment widths are scaled to fill each bar. Out of decisions (OOD) means the agent exhausted its 200-decision budget without a valid physical submission. Each accepted model response uses one decision, including invalid tool calls; this is not a token limit.
WORKLOAD & OCCLUSION
On the four tasks that vary both factors, the pooled success rate decreases from 14.4% at low workload to 2.9% at high workload. Occlusion effects are smaller and less certain.
Equal-weight task means pooled over Astra and Fable.
The same four tasks and two models.
Low → high workload: −11.5 percentage points, 95% bootstrap interval [−16.8, −6.2]. Visible → Look: −3.3 [−8.8, 1.9]; Visible → Uncover: −4.0 [−9.3, 1.2]. The occlusion intervals include zero.
SUBMISSION & STOPPING
The robot must physically press SUBMIT. A declared stop, a text answer or running out of decisions ends the episode without completing that requirement.
04 / DEMONSTRATION DATASET
Successful scripted trajectories include the physical search, inspection and testing steps needed to gather task-relevant evidence.
Each frame records three camera views, robot state, action and the current subtask. Episodes average 7.6 minutes, or about 9,100 control steps. Only successful episodes that pass recording checks are retained.
The scripted oracles can read the private specification. Their routes deliberately include information-gathering steps; these are demonstrations of investigative behavior, not evidence that the oracle itself resolves uncertainty.
Training and evaluation instances do not overlap. Evaluation holds out kitchen styles and selected task configurations, including stamp targets and bolt-chain configurations.
CONTINUE EXPLORING
Results, cost, task parameters, training, progress and failure diagnostics.
Agent interface, evaluation harness, progress equations and analysis protocol.
Exact per-task success and progress values in JSON.
Paper, code and dataset release links will be added when available.