RoboQuestDeCLaRe Lab

RoboQuest

Generalist physical agents that
search, inspect and test.

Robots must acquire task-relevant evidence through physical interaction,
then decide when they know enough to commit.

DeCLaRe Lab · Goal-directed embodied exploration

Research overview · narrated · simulation demonstrationsRead transcript ↗
01 / SearchSimulation

Where is the target?

Open storage places, track what has been checked, and retrieve the complete target set.

Task demonstration

Left scene camera · right scene camera · wrist camera

Selected demonstrations with all three camera views. Left scene camera · right scene camera · wrist camera. Aggregate results are reported separately.

10tasks across three families
3,211successful scripted demonstrations
405.5 hrecorded interaction at 20 Hz
18evaluation kitchen layouts

01 / THE BENCHMARK

Act to learn.
Then act on what you learn.

RoboQuest places a mobile manipulator in unfamiliar kitchens. The goal is given; task-critical information must be physically acquired.

01

Search

Choose where to look. A negative finding redirects the search; a discovered token opens new possibilities.

Location uncertainty
02

Inspect

Lift, turn or open an object. Decide when the evidence is sufficient to identify a hidden property.

Hidden object properties
03

Test

Intervene and observe the outcome. Infer weight, a stamp pattern, a release order or an unseen tilt.

Unknown physical behavior
SUBMIT

The robot decides when it knows enough.

A physical button press commits the result and freezes the score. Every goal condition must hold. Stopping without pressing it is a failure.

Scoring protocol ↗

Ten tasks. No prescribed investigation.

01Search

Search Room

A set of target items is stored in open or closed locations around a room, among distractor items of similar kind. The goal is to collect the complete set on the tray and press SUBMIT. The locations, the number of storage places, and which storage places are empty are not given. The robot must choose where to search, use negative findings to redirect, remember what it has checked, and stop once the set is complete. Instances vary the set size, item category, room layout, number and type of storage units, and distractors.

02Search

Locked Storage

Storage compartments open only when a matching access token rests on their reader. Tokens are either outside the initial observation or hidden inside other storage, and some compartments contain tokens for other compartments, forming a dependency chain. The goal is to retrieve the target item. Finding a token changes which places can be searched next. Closing a compartment with a token still inside locks that token away permanently, so the robot must plan token handling as well as search order. Instances vary the number of locked compartments, chain structure, token hiding places, and dead ends.

03Search

Blackout Search

A search task is performed in an unlit room. Only a light source with a glow marker and the SUBMIT button are visible at the start. The robot must acquire the light source, carry it, and search using a wrist camera that sees only what is illuminated. Because one gripper holds the light, retrieving a target requires putting the light down while aiming it at the target, then acting from memory in the dark. Instances vary the target, room, location of the light source, whether a wall switch exists, and whether a timed light turns off after a delay.

04Inspect

Marked Mugs

Several identical vessels, all mugs or all bowls in a given instance, each hold loose contents, and a label on the underside of each vessel specifies its destination. The goal is to place every vessel upright at its destination with its original contents inside. Reading a label requires lifting or tilting the vessel, which spills the contents; spilled contents stay nearby and must be returned to the correct vessel before submission. Instances vary the number of vessels, contents, label-to-destination mapping, and initial poses.

05Inspect

Painted Cubes

A collection of similar objects each carries a hidden property distributed over its faces, such as the number of painted faces, with some faces turned toward the table or away from every camera. The goal is to collect every object satisfying a stated property, for example, exactly one painted face. Partial views can reject an object but cannot confirm it, so the robot must decide how much inspection each object needs. Instances vary the number of objects, property rule, initial orientations, and whether the goal asks for one qualifying object or all of them.

06Inspect

Unfamiliar Containers

Several closed containers each open by a different mechanism, such as a hinged lid, sliding lid, drawer, or latched flip-top, and some cannot be opened at all. The target is inside one of them. The robot knows neither where the target is nor how each container opens and must discover each mechanism by interaction. Instances vary the number of containers, mix of mechanisms, target location, and number of decoys.

07Test

Stamp Composition

Several stamps with unmarked housings each carry a hidden pattern and start at an unknown rotation. A target pattern is shown on a reference. A test surface allows trial impressions, and a final surface must end with exactly the target. The robot must test stamps, read the results, select a subset and their rotations, compose the target, and press SUBMIT. Marks on the final surface are permanent. Instances vary the number of stamps, patterns, target, initial rotations, and size of the test surface.

08Test

Odd Parcel

A set of visually identical sealed parcels contains one or two parcels that differ in weight. The only instrument is a two-pan balance whose pans hold several parcels each. The goal is to place the odd parcel alone on the tray. Group weighing identify it in few comparisons, while one-by-one weighing is legal but slower. Instances vary the number of parcels, whether the odd parcel is known to be heavier or could be lighter, pan capacity, and parcel appearance.

09Test

Puzzle Box

A container is held closed by a chain of interlocking parts, where each part blocks another until it is moved. All parts are visible with large handles, but the blocking relations are not. The goal is to place the item inside on the tray. The robot must discover the release order by trying moves and observing what moves, and some wrong moves jam another part until reversed. Instances vary the number of parts, chain structure, presence of red-herring parts, and trap moves.

10Test

Wobbly Stand

A stand or table has an uneven support, tilting its surface by an amount too small to see directly. A ball placed on the surface rolls toward the low side and reveals the tilt. Shim blocks of different thicknesses are available. The goal is a level surface on which the ball stays put, with everything released. Instances vary the number of legs affected, tilt magnitude, shim set, and whether the tilt is along one axis or two.

See the ten task trajectories +
Five frames from a scripted demonstration of each of the ten RoboQuest tasks: initial scene, information gathering and task completion.
Five frames per task, from the initial scene to information gathering and completion. Open the image for a larger view.

A SHARED PHYSICAL WORLD

Mobile manipulation.
Partial information.
Consequential actions.

Environment
RoboCasa365 kitchens simulated in MuJoCo; a Franka Panda arm on a mobile base.
Observation
Two scene cameras, one wrist camera and robot proprioception. Hidden task specifications are inaccessible to the agent.
Occlusion
Visible: objects start in view. Look: move to find them. Uncover: physically remove a cover.
Consequences
Bin captures, locked-away tokens and final stamp marks can be irreversible.

THE CAPABILITIES BEHIND THE TASKS

Gather the evidence.
Use it to decide.

Ten tasks probe nine capabilities across four groups. The matrix marks direct targets and supporting requirements; success alone does not isolate any one capability.

Evidence acquisition

Directed search chooses what to investigate.
Active perception makes hidden evidence visible.

Evidence use

Evidence integration combines observations.
Memory retains what is no longer in view.

Interactive inference

Affordance discovery finds how objects work.
Experimental identification distinguishes hypotheses.
Physical causal inference relates actions to outcomes.

Action organization

Consequential action accounts for irreversible effects.
Long-horizon planning orders dependent steps.

Capabilities required by each task

Open table ↗
Task capabilities and failure taxonomy. ✓ marks a directly targeted capability, ∼ a supporting requirement, and – no designated requirement. DS: directed search. AP: active perception. EI: evidence integration. M: memory. AD: affordance discovery. XI: experimental identification. CI: physical causal inference. CA: consequential action. LP: long-horizon planning. Links to failures are shown in the failure taxonomy.
Evidence · acquisitionEvidence · useInteractive · inferenceAction · organization
FamilyTaskDSAPEIMADXICICALP
SearchSearch Room✓∼✓✓––––∼
Locked Storage✓–✓∼––∼✓✓
Blackout Search✓✓✓✓––––∼
InspectPainted Cubes–✓✓∼–––∼∼
Marked Mugs–✓∼∼–––✓∼
Unfamiliar Containers∼––∼✓∼––∼
TestStamp Composition––✓∼–✓–✓✓
Odd Parcel––✓✓–✓––∼
Puzzle Box–––∼∼∼✓∼✓
Wobbly Stand––∼∼–✓✓–∼

✓ Directly targeted · ∼ Supporting requirement · — No designated requirement. Capability links describe the task design, not independent causal measurements.

Benchmark comparison

Open table ↗
Comparison with representative manipulation benchmarks. ✓: explicit target; ∼: partial; –: not targeted. Hidden: task information absent from the current observation.
Benchmark featuresEvidence seeking
BenchmarkFocusLong hor.HiddenMemoryMobileSearchInspectTest
RLBenchVisuomotor skills∼∼–––––
ManiSkill3Scalable manipulation∼––––––
RoboTwin 2.0Bimanual manipulation∼––––∼–
LIBEROLifelong skill transfer✓––––––
CALVINLong-horizon manipulation✓∼–––––
BEHAVIOR-1KHousehold activities✓∼∼✓✓––
RoboCasa365Manipulation in kitchen✓–✓✓∼∼∼
RMBenchMemory-dependent manipulation✓✓✓–––✓
RoboQuestEvidence-seeking manipulation✓✓✓✓✓✓✓

Feature designations reproduce the manuscript comparison; they describe designated benchmark targets, not the limits of every task in each suite.

02 / EVALUATION

Progress is not
completion.

All three frontier agents play the same 50 instances per task. Even the strongest reaches 19.4% success overall, despite 44.9% mean task progress.

Equal-weight mean over 10 tasks · 500 episodes per model · percentages.

Success rate by task

0–100%
GPT-6 AstraClaude Fable 5.1Gemini 3.8 Flash
96%

Puzzle Box is the outlier.

GPT-6 Astra reaches 96% success, compared with 70% for Claude Fable 5.1 and 12% for Gemini 3.8 Flash. Excluding this task lowers their means to 10.9%, 3.3% and 0.9%.

0/150

Blackout Search remains unsolved.

None of the three models succeeds in its 50 evaluation episodes. Finding and positioning a light is only the beginning of the investigation.

Success and task progress

Open table ↗
Main results. SR: success rate. Prog: mean progress. Both in %, 50 episodes per cell. Best SR per row in bold. Overall w/o Puzzle Box excludes the easiest task, an outlier on which all three models score far above their other tasks.
GPT-6 AstraClaude Fable 5.1Gemini 3.8 Flash
FamilyTaskSRProgSRProgSRProg
SearchLocked Storage28.052.98.021.70.09.2
Search Room0.031.00.019.22.017.5
Blackout Search0.023.70.017.10.04.4
InspectPainted Cubes30.074.814.061.52.011.3
Marked Mugs10.043.00.016.60.00.0
Unfamiliar Containers6.032.20.017.00.06.3
TestPuzzle Box96.099.670.081.012.033.0
Stamp Composition14.055.32.036.90.03.4
Wobbly Stand10.018.66.014.42.04.6
Odd Parcel0.018.30.07.22.02.0
Overall19.444.910.029.32.09.2
Overall w/o Puzzle Box10.938.93.323.60.96.5

FRONTIER AGENTS

One decision. One physical command.

GPT-6 Astra and Claude Fable 5.1 use medium thinking; Gemini 3.8 Flash uses high thinking. Each decision receives three 512 × 512 images and proprioception.

Up to 200 decisions; up to 10 simulated seconds per command. The full text history is retained, with images from only the two latest observations.

TRAINED VLA BASELINE

Task-specific training falls short.

RoboQuest tests generalist physical agents: a new task should not require a new fine-tuning run. π₀.₅ is evaluated with task-specific demonstration fine-tuning, yet the reported policy still falls short of the investigation and completion capabilities these tasks require.

The reported five-task evaluation reaches 1.9% success and 42.7% progress on Puzzle Box. Marked Mugs, Wobbly Stand, Stamp Composition and Odd Parcel have 0% success.

These results expose a capability gap in the evaluated π₀.₅ setup, even with task-specific training. They do not establish that all VLA architectures are incapable of these tasks.

Training and evaluation details ↗

Cost and interaction effort

Open table ↗
Cost and effort per episode. Sim / Wall: simulated / wall-clock minutes. In / Out: tokens in millions. Costs in USD at list prices.
Effort per episodeTokens per episodeCost (USD)
ModelDecisionsSim min.Wall min.In (M)Out (M)Cached (%)$/episode$/success
GPT-6 Astra143.710.453.85.820.0291.811.3058.2
Claude Fable 5.1148.86.154.714.550.1194.219.22192.2
Gemini 3.8 Flash160.85.141.019.660.1389.33.37168.3

Costs are the manuscript’s recorded USD list-price accounting, not a current pricing quotation. Cost per success divides total spend by successful episodes.

03 / FAILURE ANALYSIS

Where does
exploration break down?

A failed episode can leave several requirements unmet. The analysis attributes each unmet requirement to the first uncorrected failure in the interaction sequence.

FAILURE TAXONOMY

Seven ways the
investigation breaks.

Failures are organized by where the connection between evidence, decision and physical action first breaks and remains uncorrected.

Before action

Acquire → Interpret → Plan

Was the necessary evidence gathered, understood and used?

During action

Execute

Did the command achieve the intended physical result?

After feedback

Detect → Recover → Stop

Was the outcome checked, repaired if needed, and completion judged correctly?

Definitions and examples from the failure taxonomy. Capability links follow the manuscript’s category-level mapping.
PhaseFailure categoryWhat breaksExampleLinked capabilities
Before actionInadequate evidence gatheringNeeded evidence was never acquired.A location remains unchecked, a hidden face unobserved, or a diagnostic test omitted.Evidence acquisition · Interactive inference
Before actionMisinterpretation of evidenceThe evidence was observed but read incorrectly.A label is misread or a trial outcome assigned the wrong meaning.Evidence use
Before actionFaulty reasoning or planningAvailable evidence leads to an invalid decision or plan.Balance comparisons are combined incorrectly, or a dependency is ignored.Evidence use · Interactive inference · Action organization
During actionFailed action executionThe intended physical outcome is not achieved.A grasp misses, a carried object falls, or a shim lands beside its leg.Action organization
After feedbackFailure to detect errorsA consequential error is not recognized.A dropped object is left behind as the agent continues.Evidence use · Action organization
After feedbackIneffective error recoveryThe agent attempts a repair but does not correct the error.Repeated retrieval attempts leave the fallen object on the floor.Interactive inference · Action organization
After feedbackIncorrect terminationThe agent submits or stops with reachable goals unfinished.A vessel is never handled before the robot ends the episode.Evidence use · Action organization

INADEQUATE EVIDENCE GATHERING

The largest failure category
comes before execution.

36.7%GPT-6 Astra
42.2%Claude Fable 5.1

Share of unmet requirements in failed episodes. Denominators: 901 for Astra and 1,129 for Fable; 95% bootstrap intervals: [34, 40] and [39, 45].

01 / Evidence acquisition

Search ends too soon.

In Search Room and Blackout Search, the gripper never approached within 25 cm of 83% of Astra’s missed targets and 91% of Fable’s. Relevant places remain unchecked.

02 / Evidence use

More items, less success.

From low to high workload, success falls by 15.8 percentage points for Astra and 7.2 for Fable on the four tasks that vary both workload and occlusion.

03 / Interactive inference

Hidden locks separate the agents.

On the covered Puzzle Box subset, Astra reaches 96.4% success. Fable drops from 100% without the cover to 34.6% with it.

04 / Action organization

Evidence still has to become action.

Failed execution accounts for 21.1% of Astra’s unmet requirements and 15.4% of Fable’s. Incorrect termination contributes 11.3% and 20.5%.

A separate analysis cohort. 540 matched instances per model, with 46–63 per task. This differs from the 500-episode main evaluation. The breakdown contains 432 failed episodes for Astra and 486 for Fable; six ambiguous Odd Parcel outcomes are excluded. These percentages are descriptive attributions, not isolated causal effects.

Failure breakdown by category

Open table ↗
Failure breakdown by taxonomy category, in % of the unmet requirements of failed episodes. Overall: 95% bootstrap intervals. the attribution protocol gives the attribution rules.
SearchInspectTestOverall
PhaseCategoryLinked capabilitiesAstraFableAstraFableAstraFableAstraFable
Before actionInadequate evidence gatheringAcquisition, inference50.659.739.040.915.225.236.7 · [34, 40]42.2 · [39, 45]
Misinterpretation of evidenceUse16.620.53.92.415.26.211.5 · [9, 14]9.0 · [7, 11]
Faulty reasoning or planningUse, inference, organization0.60.34.84.420.115.77.4 · [6, 9]6.4 · [5, 8]
During actionFailed action executionOrganization7.44.628.117.329.924.321.1 · [19, 23]15.4 · [13, 17]
After feedbackFailure to detect errorsUse, organization14.16.33.31.30.00.06.3 · [5, 8]2.5 · [2, 3]
Ineffective error recoveryInference, organization10.78.63.03.10.00.05.0 · [4, 6]3.9 · [3, 5]
Incorrect terminationUse, organization0.00.016.330.419.728.611.3 · [8, 15]20.5 · [17, 24]
Failed episodes161173137155134158432486
Unmet requirements (denominator)3263473314572443259011129

TASK-LEVEL DIAGNOSIS

The failure mix changes
with the task.

Search failures are dominated by missing evidence. Marked Mugs and Wobbly Stand more often break during execution. Inspect each task’s distribution below.

Shares of unmet requirements within failed episodes in the matched analysis cohort. Displayed percentages are the reported rounded values; segment widths are scaled to fill each bar. Out of decisions (OOD) means the agent exhausted its 200-decision budget without a valid physical submission. Each accepted model response uses one decision, including invalid tool calls; this is not a token limit.

Complete failure breakdown by task +

Failure categories by task

Open table ↗
Failure breakdown per task. Share (%) of unmet requirements in each category. Ep.: failed episodes; Req.: unmet requirements. IEG: inadequate evidence gathering; MIS: misinterpretation of evidence; FRP: faulty reasoning or planning; FAE: failed action execution; FDE: failure to detect errors; IER: ineffective error recovery; IT: incorrect termination; OOD: out of decisions.
TaskPolicyEp.Req.IEGMISFRPFAEFDEIERITOOD
Search RoomGPT-6 Astra6315059150613800
Claude Fable 5.1631606816065400
Locked StorageGPT-6 Astra4345242940222000
Claude Fable 5.1555824402092600
Blackout SearchGPT-6 Astra551315015011131100
Claude Fable 5.1551296517057600
Painted CubesGPT-6 Astra378558150410220
Claude Fable 5.1471326170012300
Marked MugsGPT-6 Astra491282007024221
Claude Fable 5.1541874104213491
Unfamiliar ContainersGPT-6 Astra511186601416463
Claude Fable 5.1541387201413460
Stamp CompositionGPT-6 Astra4212352703100370
Claude Fable 5.151170151102400490
Odd ParcelGPT-6 Astra46754105900000
Claude Fable 5.146946004000000
Puzzle BoxGPT-6 Astra220001000000
Claude Fable 5.115150053470000
Wobbly StandGPT-6 Astra44440911750050
Claude Fable 5.1464602116700200

WORKLOAD & OCCLUSION

More to keep track of.
Less task success.

On the four tasks that vary both factors, the pooled success rate decreases from 14.4% at low workload to 2.9% at high workload. Occlusion effects are smaller and less certain.

Workload

Equal-weight task means pooled over Astra and Fable.

Occlusion

The same four tasks and two models.

Low → high workload: −11.5 percentage points, 95% bootstrap interval [−16.8, −6.2]. Visible → Look: −3.3 [−8.8, 1.9]; Visible → Uncover: −4.0 [−9.3, 1.2]. The occlusion intervals include zero.

Success, progress and confidence intervals +

Workload and occlusion contrasts

Open table ↗
Workload and occlusion on the four tasks that vary both, pooled over both models. Brackets: 95% bootstrap intervals. Last block: workload drop minus occlusion drop.
SRProg
AstraFablePooledAstraFablePooled
Workload
Low21.77.214.453.031.342.1
Mid12.84.48.646.833.039.9
High5.80.02.941.427.134.2
High - Low-15.8 · [-25.6, -6.7]-7.2 · [-11.7, -2.8]-11.5 · [-16.8, -6.2]-11.6 · [-20.5, -2.8]-4.2 · [-11.7, 3.3]-7.9 · [-14.0, -1.8]
Occlusion
Visible16.75.611.151.134.242.7
Look10.84.77.845.627.636.6
Uncover12.81.47.144.529.537.0
Look - Visible-5.8 · [-15.0, 3.3]-0.8 · [-5.6, 5.3]-3.3 · [-8.8, 1.9]-5.6 · [-14.3, 3.5]-6.6 · [-13.8, 0.7]-6.1 · [-11.6, -0.4]
Uncover - Visible-3.9 · [-13.9, 5.6]-4.2 · [-8.3, 0.0]-4.0 · [-9.3, 1.2]-6.6 · [-15.2, 1.8]-4.7 · [-11.9, 2.7]-5.6 · [-11.6, 0.4]
Workload drop minus occlusion drop
vs. Look10.0 · [-3.1, 23.1]6.4 · [1.4, 12.5]8.2 · [1.5, 15.0]6.1 · [-6.0, 18.9]-2.4 · [-12.0, 7.6]1.8 · [-6.3, 9.8]
vs. Uncover11.9 · [-0.6, 24.2]3.1 · [0.0, 7.5]7.5 · [1.1, 13.8]5.0 · [-6.9, 16.9]-0.5 · [-10.1, 9.3]2.3 · [-5.9, 10.5]

INTERACTIVE INFERENCE

When the locks disappear,
the agents diverge.

Covering the mechanism removes a visual shortcut. Astra maintains high success, while Fable’s success falls sharply. These comparisons use Puzzle Box chain lengths four and five.

Hidden-lock results and uncertainty +

Puzzle Box with hidden locks

Open table ↗
Puzzle Box: effect of hiding the locks, on chain lengths 4 and 5. Δ: Covered minus No cover, with 95% bootstrap interval. Last row: Claude Fable 5.1's Δ minus GPT-6 Astra's.
PolicyCoverSRProgDecisions
GPT-6 AstraNo cover94.499.684.7
Covered96.499.5119.5
Δ2.0 · [-8.7, 16.7]-0.1 · [-1.4, 1.0]34.8 · [14.0, 55.8]
Claude Fable 5.1No cover100.0100.0106.7
Covered34.659.1174.6
Δ-65.4 · [-83.7, -46.2]-40.9 · [-56.3, -26.0]67.9 · [46.8, 86.1]
Gemini 3.8 FlashNo cover27.845.2172.2
Covered0.025.5199.8
Δ-27.8 · [-44.4, -11.1]-19.7 · [-37.9, -1.0]27.6 · [8.5, 48.8]
Difference in Δ-67.3 · [-89.8, -44.9]-40.8 · [-56.2, -26.1]33.1 · [5.0, 63.6]

SUBMISSION & STOPPING

Reaching a state
is not a valid submission.

The robot must physically press SUBMIT. A declared stop, a text answer or running out of decisions ends the episode without completing that requirement.

How episodes end

Open table ↗
How episodes end, in % of 500 episodes per model. Stop, Text answer, and Timeout end without a SUBMIT and all fail. Precision: share of submissions that were correct.
SubmittedNo submit
ModelCorrectWrongStoppedText answerTimeoutSubmit ratePrecision
GPT-6 Astra19.465.65.40.09.685.022.8
Claude Fable 5.110.055.028.65.41.065.015.4
Gemini 3.8 Flash2.055.60.80.041.657.63.5
Episode outcomes for every task +

Episode outcomes by task

Open table ↗
Episode outcomes per task. Percentage of episodes by outcome, with absolute counts shown below in parentheses. Cell shading indicates the percentage within each model–task pair (darker indicates a larger percentage). C: correct submit. W: wrong submit. S: declared stop. A: text answer instead of a submit. T: timeout (out of decisions). Tasks are ordered as in Success and task progress.
GPT-6 AstraClaude Fable 5.1Gemini 3.8 Flash
FamilyTaskCWSATCWSATCWSAT
SearchLocked Storage28%(14)44%(22)20%(10)0%(0)8%(4)8%(4)42%(21)36%(18)14%(7)0%(0)0%(0)20%(10)2%(1)0%(0)78%(39)
Search Room0%(0)56%(28)8%(4)0%(0)36%(18)0%(0)56%(28)34%(17)8%(4)2%(1)2%(1)32%(16)0%(0)0%(0)66%(33)
Blackout Search0%(0)76%(38)12%(6)0%(0)12%(6)0%(0)46%(23)50%(25)2%(1)2%(1)0%(0)56%(28)6%(3)0%(0)38%(19)
InspectPainted Cubes30%(15)70%(35)0%(0)0%(0)0%(0)14%(7)84%(42)2%(1)0%(0)0%(0)2%(1)82%(41)0%(0)0%(0)16%(8)
Marked Mugs10%(5)82%(41)4%(2)0%(0)4%(2)0%(0)80%(40)20%(10)0%(0)0%(0)0%(0)88%(44)0%(0)0%(0)12%(6)
Unfamiliar Containers6%(3)66%(33)6%(3)0%(0)22%(11)0%(0)24%(12)68%(34)8%(4)0%(0)0%(0)36%(18)0%(0)0%(0)64%(32)
TestPuzzle Box96%(48)4%(2)0%(0)0%(0)0%(0)70%(35)0%(0)14%(7)16%(8)0%(0)12%(6)6%(3)0%(0)0%(0)82%(41)
Stamp Composition14%(7)84%(42)2%(1)0%(0)0%(0)2%(1)74%(37)22%(11)2%(1)0%(0)0%(0)54%(27)0%(0)0%(0)46%(23)
Wobbly Stand10%(5)86%(43)2%(1)0%(0)2%(1)6%(3)62%(31)30%(15)0%(0)2%(1)2%(1)96%(48)0%(0)0%(0)2%(1)
Odd Parcel0%(0)88%(44)0%(0)0%(0)12%(6)0%(0)82%(41)10%(5)4%(2)4%(2)2%(1)86%(43)0%(0)0%(0)12%(6)
All tasks19.4%(97)65.6%(328)5.4%(27)0.0%(0)9.6%(48)10.0%(50)55.0%(275)28.6%(143)5.4%(27)1.0%(5)2.0%(10)55.6%(278)0.8%(4)0.0%(0)41.6%(208)

04 / DEMONSTRATION DATASET

Learning to investigate.

Successful scripted trajectories include the physical search, inspection and testing steps needed to gather task-relevant evidence.

3,211successful episodes
405.5hours of interaction
20 Hzcamera, state and action records

Each frame records three camera views, robot state, action and the current subtask. Episodes average 7.6 minutes, or about 9,100 control steps. Only successful episodes that pass recording checks are retained.

The scripted oracles can read the private specification. Their routes deliberately include information-gathering steps; these are demonstrations of investigative behavior, not evidence that the oracle itself resolves uncertainty.

Training and evaluation instances do not overlap. Evaluation holds out kitchen styles and selected task configurations, including stamp targets and bolt-chain configurations.

Demonstration dataset

Open table ↗
Demonstration dataset. Scripted information gathering in each oracle's route, which the oracle performs although it knows the hidden information. Episodes: successful demonstrations. Hours: interaction time at 20 Hz.
TaskScripted information gatheringEpisodesHours
Locked Storagefollows the token chain27833.2
Search Roomopens compartments nearest-first until the targets are seen29446.8
Blackout Searchcarries the lamp along a nearest-first search20153.0
Painted Cubesremoves covers and turns each cube face by face27650.3
Marked Mugslifts and tilts each vessel to read its label31220.9
Unfamiliar Containerssome episodes make one or two wrong tries before opening39876.9
Puzzle Boxtries sliders, one to three of which are blocked44721.1
Stamp Compositionprints test impressions before the final one36332.6
Wobbly Standreleases the ball, places a shim, and tests again38118.9
Odd Parcelweighs parcels on the balance26151.7
Total3,211405.5

CONTINUE EXPLORING

The full experimental picture.

Paper, code and dataset release links will be added when available.