RoboQuestDeCLaRe Lab

THE COMPLETE RESULTS & PROTOCOL

Research tables.

All 24 tables from the current manuscript, including the full diagnostic analysis. Values, confidence intervals and denominators are retained. Tables are organized by topic.

24 tables

Benchmark comparison

#
Comparison with representative manipulation benchmarks. ✓: explicit target; ∼: partial; –: not targeted. Hidden: task information absent from the current observation.
Benchmark featuresEvidence seeking
BenchmarkFocusLong hor.HiddenMemoryMobileSearchInspectTest
RLBenchVisuomotor skills∼∼–––––
ManiSkill3Scalable manipulation∼––––––
RoboTwin 2.0Bimanual manipulation∼––––∼–
LIBEROLifelong skill transfer✓––––––
CALVINLong-horizon manipulation✓∼–––––
BEHAVIOR-1KHousehold activities✓∼∼✓✓––
RoboCasa365Manipulation in kitchen✓–✓✓∼∼∼
RMBenchMemory-dependent manipulation✓✓✓–––✓
RoboQuestEvidence-seeking manipulation✓✓✓✓✓✓✓

Capabilities required by each task

#
Task capabilities and failure taxonomy. ✓ marks a directly targeted capability, ∼ a supporting requirement, and – no designated requirement. DS: directed search. AP: active perception. EI: evidence integration. M: memory. AD: affordance discovery. XI: experimental identification. CI: physical causal inference. CA: consequential action. LP: long-horizon planning. Links to failures are shown in the failure taxonomy.
Evidence · acquisitionEvidence · useInteractive · inferenceAction · organization
FamilyTaskDSAPEIMADXICICALP
SearchSearch Room✓∼✓✓––––∼
Locked Storage✓–✓∼––∼✓✓
Blackout Search✓✓✓✓––––∼
InspectPainted Cubes–✓✓∼–––∼∼
Marked Mugs–✓∼∼–––✓∼
Unfamiliar Containers∼––∼✓∼––∼
TestStamp Composition––✓∼–✓–✓✓
Odd Parcel––✓✓–✓––∼
Puzzle Box–––∼∼∼✓∼✓
Wobbly Stand––∼∼–✓✓–∼

Success and task progress

#
Main results. SR: success rate. Prog: mean progress. Both in %, 50 episodes per cell. Best SR per row in bold. Overall w/o Puzzle Box excludes the easiest task, an outlier on which all three models score far above their other tasks.
GPT-6 AstraClaude Fable 5.1Gemini 3.8 Flash
FamilyTaskSRProgSRProgSRProg
SearchLocked Storage28.052.98.021.70.09.2
Search Room0.031.00.019.22.017.5
Blackout Search0.023.70.017.10.04.4
InspectPainted Cubes30.074.814.061.52.011.3
Marked Mugs10.043.00.016.60.00.0
Unfamiliar Containers6.032.20.017.00.06.3
TestPuzzle Box96.099.670.081.012.033.0
Stamp Composition14.055.32.036.90.03.4
Wobbly Stand10.018.66.014.42.04.6
Odd Parcel0.018.30.07.22.02.0
Overall19.444.910.029.32.09.2
Overall w/o Puzzle Box10.938.93.323.60.96.5

Cost and interaction effort

#
Cost and effort per episode. Sim / Wall: simulated / wall-clock minutes. In / Out: tokens in millions. Costs in USD at list prices.
Effort per episodeTokens per episodeCost (USD)
ModelDecisionsSim min.Wall min.In (M)Out (M)Cached (%)$/episode$/success
GPT-6 Astra143.710.453.85.820.0291.811.3058.2
Claude Fable 5.1148.86.154.714.550.1194.219.22192.2
Gemini 3.8 Flash160.85.141.019.660.1389.33.37168.3

Action tools

#
Action tools. Positions are in world metres; a tick is 0.05 s of simulated time. Every tool except stop accepts a tick count from 1 to 200.
ToolArgumentsDefault ticks
armgripper-site position (required); orientation as a quaternion; gripper command in [-1,1], where -1 opens and +1 closes120
basebody-frame forward, left, and yaw velocity, each in [-0.5, 0.5]20
waitoptional gripper command; holds the arm and base20
stopnone; abandons the episode, which fails–

Training and inference settings

#
π₀.₅ Training and Inference Hyperparameters.
CategoryHyperparameterValue
TrainingOptimizerAdamW (β₁=0.9, β₂=0.999, ε=10⁻⁸)
Weight Decay1× 10⁻⁴
Peak Learning Rate2.5× 10⁻⁵
Learning Rate ScheduleCosine decay with 1000 warmup steps
Global Batch Size32
Gradient Clipping Norm1.0
Flow Loss Multiplier10.0
Action NormalizationQuantile (q₀₁, q₉₉)
InferenceReplanning FrequencyOnce per 20 ticks (1.0 s)
Flow Denoising Steps10

Task parameters

#
Tasks and their main parameters. Occlusion: Visible (V), Look (L), Uncover (U), defined in the environment description. W, O, C: tasks used for analysis of workload, occlusion, and cover in the diagnostic analysis.
TaskHidden informationMain parametersWOC
Search
Locked Storagetarget behind lockschain depth 1/2/3✓
Search Roomwhere the targets arecompartments 4/6/8targets 2/3
Blackout Searchtargets in the darkcompartments 4/6/8targets 2/3
Inspect
Painted Cubesmarks on unseen facescubes 4/6/8occlusion V/L/U✓✓
Marked Mugslabels under vesselsvessels 3/4/5occlusion V/L/U✓✓
Unfamiliar Containershow containers openitems 2/3/4occlusion V/L✓
Test
Puzzle Boxbolt blocking orderchain 4/5/6cover no/yes✓
Stamp Compositionpatterns and rotationsstamps 2/3/4occlusion V/L/U✓✓
Wobbly Standshort legs and gapsshort legs 1/2shim choice no/yes
Odd Parcelwhich parcel is oddparcels 4/6/8occlusion V/L/U✓✓

Task progress definitions

#
Progress P at a glance. Every stage counts only in a passive state (released, supported, untouched where the predicate requires it); P is uncapped, frozen at the first Submit press and clipped to [0,1].
TaskUnitsStages or termsDeduction δ
Search roomtargetsopened / out / placednon-target on tray
Locked storagechain links, targetlinks + opened / out / placednon-target on tray
Blackout searchtargetsopened / out / placednon-target on tray
Marked mugsvessels{0,0.5,1} per vessel–
Painted cubescubescorrect captures–
Unfamiliar containersitemsopened / out / in bowl–
Stamp compositiontarget cellscell grades {0,0.5,1}, strays in the denominator–
Odd parcelparcel setJaccard of boxed set and odd setpans loaded
Puzzle boxchain bolts, itembolts + lid / out / tray–
Wobbly standstandtilt removed q, ball on top b–

Episode outcomes by task

#
Episode outcomes per task. Percentage of episodes by outcome, with absolute counts shown below in parentheses. Cell shading indicates the percentage within each model–task pair (darker indicates a larger percentage). C: correct submit. W: wrong submit. S: declared stop. A: text answer instead of a submit. T: timeout (out of decisions). Tasks are ordered as in Success and task progress.
GPT-6 AstraClaude Fable 5.1Gemini 3.8 Flash
FamilyTaskCWSATCWSATCWSAT
SearchLocked Storage28%(14)44%(22)20%(10)0%(0)8%(4)8%(4)42%(21)36%(18)14%(7)0%(0)0%(0)20%(10)2%(1)0%(0)78%(39)
Search Room0%(0)56%(28)8%(4)0%(0)36%(18)0%(0)56%(28)34%(17)8%(4)2%(1)2%(1)32%(16)0%(0)0%(0)66%(33)
Blackout Search0%(0)76%(38)12%(6)0%(0)12%(6)0%(0)46%(23)50%(25)2%(1)2%(1)0%(0)56%(28)6%(3)0%(0)38%(19)
InspectPainted Cubes30%(15)70%(35)0%(0)0%(0)0%(0)14%(7)84%(42)2%(1)0%(0)0%(0)2%(1)82%(41)0%(0)0%(0)16%(8)
Marked Mugs10%(5)82%(41)4%(2)0%(0)4%(2)0%(0)80%(40)20%(10)0%(0)0%(0)0%(0)88%(44)0%(0)0%(0)12%(6)
Unfamiliar Containers6%(3)66%(33)6%(3)0%(0)22%(11)0%(0)24%(12)68%(34)8%(4)0%(0)0%(0)36%(18)0%(0)0%(0)64%(32)
TestPuzzle Box96%(48)4%(2)0%(0)0%(0)0%(0)70%(35)0%(0)14%(7)16%(8)0%(0)12%(6)6%(3)0%(0)0%(0)82%(41)
Stamp Composition14%(7)84%(42)2%(1)0%(0)0%(0)2%(1)74%(37)22%(11)2%(1)0%(0)0%(0)54%(27)0%(0)0%(0)46%(23)
Wobbly Stand10%(5)86%(43)2%(1)0%(0)2%(1)6%(3)62%(31)30%(15)0%(0)2%(1)2%(1)96%(48)0%(0)0%(0)2%(1)
Odd Parcel0%(0)88%(44)0%(0)0%(0)12%(6)0%(0)82%(41)10%(5)4%(2)4%(2)2%(1)86%(43)0%(0)0%(0)12%(6)
All tasks19.4%(97)65.6%(328)5.4%(27)0.0%(0)9.6%(48)10.0%(50)55.0%(275)28.6%(143)5.4%(27)1.0%(5)2.0%(10)55.6%(278)0.8%(4)0.0%(0)41.6%(208)

Demonstration dataset

#
Demonstration dataset. Scripted information gathering in each oracle's route, which the oracle performs although it knows the hidden information. Episodes: successful demonstrations. Hours: interaction time at 20 Hz.
TaskScripted information gatheringEpisodesHours
Locked Storagefollows the token chain27833.2
Search Roomopens compartments nearest-first until the targets are seen29446.8
Blackout Searchcarries the lamp along a nearest-first search20153.0
Painted Cubesremoves covers and turns each cube face by face27650.3
Marked Mugslifts and tilts each vessel to read its label31220.9
Unfamiliar Containerssome episodes make one or two wrong tries before opening39876.9
Puzzle Boxtries sliders, one to three of which are blocked44721.1
Stamp Compositionprints test impressions before the final one36332.6
Wobbly Standreleases the ball, places a shim, and tests again38118.9
Odd Parcelweighs parcels on the balance26151.7
Total3,211405.5

Workload and occlusion estimates

#
Factor estimates. Task-balanced aggregate success rate and progress (%) per factor level, and contrasts between levels, with cell-stratified paired bootstrap 95% confidence intervals (4,000 resamples). Workload uses the six workload tasks; occlusion uses the four occlusion tasks; the shared-task rows recompute the workload contrasts on the four occlusion tasks only, and the last block gives the direct difference between the workload penalty (Low - High) and each occlusion penalty (Visible - Look, Visible - Uncover) on those tasks.
GPT-6 AstraClaude Fable 5.1
AnalysisQuantitySRProgSRProg
Workload (6 tasks)Low23.1 [16.1, 30.3]51.1 [44.9, 57.2]9.0 [4.9, 13.2]30.9 [25.3, 36.7]
Mid12.6 [7.5, 18.2]45.7 [40.0, 51.1]3.0 [0.0, 6.1]27.0 [22.9, 31.1]
High7.7 [3.0, 12.6]39.0 [33.9, 44.1]0.0 [0.0, 0.0]22.4 [18.9, 26.0]
Mid - Low-10.5 [-19.1, -1.7]-5.3 [-13.5, 3.0]-6.0 [-11.3, -0.8]-3.9 [-10.8, 3.0]
High - Low-15.4 [-24.0, -6.8]-12.1 [-20.1, -3.9]-9.0 [-13.2, -4.9]-8.5 [-15.4, -1.9]
Occlusion (4 tasks)Visible16.7 [9.7, 23.6]51.1 [44.8, 57.5]5.6 [2.8, 8.3]34.2 [29.1, 39.3]
Look10.8 [4.7, 17.2]45.6 [39.1, 51.8]4.7 [0.0, 9.7]27.6 [22.6, 32.8]
Uncover12.8 [6.4, 19.7]44.5 [38.7, 50.6]1.4 [0.0, 4.2]29.5 [24.0, 34.8]
Look - Visible-5.8 [-15.0, 3.3]-5.6 [-14.3, 3.5]-0.8 [-5.6, 5.3]-6.6 [-13.8, 0.7]
Uncover - Visible-3.9 [-13.9, 5.6]-6.6 [-15.2, 1.8]-4.2 [-8.3, 0.0]-4.7 [-11.9, 2.7]
Workload, shared 4 tasksLow21.7 [13.6, 29.4]53.0 [46.3, 59.8]7.2 [2.8, 11.7]31.3 [25.3, 37.4]
Mid12.8 [6.9, 18.9]46.8 [41.0, 52.8]4.4 [0.0, 9.2]33.0 [27.6, 38.3]
High5.8 [1.4, 11.4]41.4 [35.4, 47.1]0.0 [0.0, 0.0]27.1 [22.8, 31.4]
Mid - Low-8.9 [-19.2, 1.4]-6.2 [-15.1, 2.6]-2.8 [-8.9, 3.6]1.7 [-6.4, 9.7]
High - Low-15.8 [-25.6, -6.7]-11.6 [-20.5, -2.8]-7.2 [-11.7, -2.8]-4.2 [-11.7, 3.3]
Difference of penaltiesWorkload - Look penalty10.0 [-3.1, 23.1]6.1 [-6.0, 18.9]6.4 [1.4, 12.5]-2.4 [-12.0, 7.6]
Workload - Uncover penalty11.9 [-0.6, 24.2]5.0 [-6.9, 16.9]3.1 [0.0, 7.5]-0.5 [-10.1, 9.3]

Occlusion by task

#
Occlusion by task. Visible / Look / Uncover on the four tasks that share the occlusion design. Visible means no additional object search or cover removal is required; the deciding information (die faces, parcel mass, bottom labels, cube faces) is still hidden in every mode. Each mode is the mean of its workload cells. SR and Prog in %, Dec: mean decisions, T/O: timeout share (%).
TaskPolicySR V / L / UProg V / L / UDec V / L / UT/O V / L / U
Stamp CompositionGPT-6 Astra6 / 31 / 762 / 63 / 39110 / 86 / 1190 / 0 / 0
Claude Fable 5.10 / 6 / 039 / 36 / 38143 / 139 / 1430 / 0 / 0
Odd ParcelGPT-6 Astra0 / 0 / 014 / 20 / 14174 / 186 / 18913 / 11 / 21
Claude Fable 5.10 / 0 / 09 / 4 / 7181 / 180 / 1887 / 0 / 5
Marked MugsGPT-6 Astra17 / 0 / 1347 / 36 / 46132 / 121 / 12611 / 0 / 0
Claude Fable 5.10 / 0 / 021 / 15 / 15133 / 139 / 1530 / 0 / 0
Painted CubesGPT-6 Astra44 / 12 / 3181 / 63 / 79114 / 122 / 1300 / 0 / 0
Claude Fable 5.122 / 13 / 668 / 56 / 59133 / 136 / 1430 / 0 / 0
Task-balanced meanGPT-6 Astra16.7 / 10.8 / 12.851.1 / 45.6 / 44.5133 / 129 / 1416.1 / 2.9 / 5.4
Claude Fable 5.15.6 / 4.7 / 1.434.2 / 27.6 / 29.5148 / 149 / 1571.7 / 0.0 / 1.2

Workload × occlusion

#
Workload × occlusion cells. Success rate / progress (%) per design cell for the four tasks that cross both factors, with the number of episodes per cell in parentheses (the same instances for both models). Rows: workload level; columns: occlusion variant.
GPT-6 AstraClaude Fable 5.1
TaskWorkloadVisibleLookUncoverVisibleLookUncover
Stamp CompositionLow (2)17 / 74 (6)60 / 84 (5)20 / 44 (5)0 / 23 (6)0 / 20 (5)0 / 47 (5)
Mid (3)0 / 47 (5)17 / 71 (6)0 / 50 (6)0 / 72 (5)17 / 44 (6)0 / 19 (6)
High (4)0 / 67 (6)17 / 35 (6)0 / 22 (5)0 / 24 (6)0 / 43 (6)0 / 47 (5)
Odd ParcelLow (4)0 / 26 (7)0 / 26 (7)0 / 0 (2)0 / 13 (7)0 / 0 (7)0 / 0 (2)
Mid (6)0 / 8 (5)0 / 24 (5)0 / 30 (7)0 / 8 (5)0 / 8 (5)0 / 9 (7)
High (8)0 / 9 (5)0 / 9 (5)0 / 13 (7)0 / 5 (5)0 / 5 (5)0 / 11 (7)
Marked MugsLow (3)17 / 53 (6)0 / 37 (5)0 / 50 (5)0 / 28 (6)0 / 27 (5)0 / 30 (5)
Mid (4)17 / 42 (6)0 / 31 (6)20 / 42 (5)0 / 27 (6)0 / 15 (6)0 / 10 (5)
High (5)17 / 47 (6)0 / 40 (6)20 / 46 (5)0 / 7 (6)0 / 3 (6)0 / 4 (5)
Painted CubesLow (4)67 / 88 (6)20 / 65 (5)60 / 90 (5)67 / 83 (6)20 / 50 (5)0 / 55 (5)
Mid (6)67 / 92 (6)0 / 50 (5)33 / 75 (6)0 / 67 (6)20 / 57 (5)17 / 61 (6)
High (8)0 / 62 (5)17 / 75 (6)0 / 71 (6)0 / 55 (5)0 / 60 (6)0 / 60 (6)

Additional workload comparisons

#
Workload cells of the two remaining workload tasks. Success rate / progress (%) per design cell with episode counts in parentheses. Unfamiliar Containers crosses item count with visible / look; Locked Storage crosses chain depth with the presence of a dead-end link.
GPT-6 AstraClaude Fable 5.1
TaskWorkloadVisibleLookVisibleLook
Unfamiliar ContainersLow (2)20 / 42 (10)9 / 39 (11)0 / 17 (10)0 / 18 (11)
Mid (3)0 / 36 (8)0 / 21 (7)0 / 26 (8)0 / 13 (7)
High (4)0 / 27 (7)0 / 19 (7)0 / 15 (7)0 / 11 (7)
TaskWorkloadNo dead endDead endNo dead endDead end
Locked StorageLow (1)38 / 55 (8)38 / 53 (8)38 / 38 (8)12 / 48 (8)
Mid (2)38 / 56 (8)11 / 62 (9)0 / 8 (8)0 / 12 (9)
High (3)33 / 58 (9)12 / 32 (8)0 / 19 (9)0 / 6 (8)

Failure attribution rules

#
Attribution rules per task. Unit: what counts as one unmet requirement. The rules are applied in the order listed, and the first matching rule assigns the category. IEG: inadequate evidence gathering; MIS: misinterpretation of evidence; FRP: faulty reasoning or planning; FAE: failed action execution; FDE/IER: after an execution failure, no repair attempt or a failed one; IT: incorrect termination.
Task and unitRules
Search Room, Blackout Search; each missing targetOn the tray but not upright: FAE. Carried and lost: FDE/IER. Hiding place never within 25 cm: IEG. Three or more grasp attempts at the target without carrying it: IER. Gripper within 15 cm of the target: MIS. Otherwise IEG. Last, a look-alike of the same colour or kind on the tray turns an IEG or MIS unit into MIS; an unpaired non-target on the tray is its own MIS unit.
Locked Storage; the targetChain progress is read off the progress score. If the target is reachable, as in Search Room. Otherwise, for the token of the first closed link: never approached, IEG; carried and lost, FDE/IER; placed on its reader but the link is closed at the press, FRP.
Unfamiliar Containers; each item not in the bowlBox never opened: IEG if never approached or probed fewer than three times, FRP if grasped three or more times. Box opened, item still inside: IER after two or more grasp attempts, else IT. Item out of its box: FDE/IER.
Painted Cubes; each misfiled or uncaptured cubeFell into a bin when the grip was lost: FAE. Wrong bin chosen: IEG if some set of unseen faces explains the choice, else MIS. Hidden and never picked: IEG. In view and never picked, or set down and never binned: IT. Carried and lost: FDE/IER.
Marked Mugs; each failed vesselWrong pad: MIS if the vessel was tilted before placement, else IEG. Correct pad with another vessel's balls: FRP. Correct pad with own balls missing: IER if balls were picked up again, else FDE. Tipped over, or set down at its pad without resting on it: FAE. Never moved: IT.
Stamp Composition; each wrong cellInk outside the grid, or dots across cell edges: FAE. Blank target cells with fewer final impressions than the solution needs: IT. Other wrong cells: IEG if a stamp reached the final board untested, else MIS.
Odd Parcel; each missing or extra parcelExtra parcel in the box: FRP. Missing odd parcel: IEG if the agent made fewer pan placements than there are parcels, else FRP.
Puzzle Box; the stallFirst unreleased part in the release order: blocked at the press, FRP; never grasped, IEG; grasped three or more times, FRP; else FAE. Lid never opened after all bolts: IEG if never grasped, else FAE. Item out but not in the tray: FAE.
Wobbly Stand; the standTilt above 15°, shim beside a leg, or level without the ball: FAE. Shim under a leg that is not short: MIS. Short legs partly shimmed or wrong thickness: FRP. No shim placed: IT.

Failure categories by task

#
Failure breakdown per task. Share (%) of unmet requirements in each category. Ep.: failed episodes; Req.: unmet requirements. IEG: inadequate evidence gathering; MIS: misinterpretation of evidence; FRP: faulty reasoning or planning; FAE: failed action execution; FDE: failure to detect errors; IER: ineffective error recovery; IT: incorrect termination; OOD: out of decisions.
TaskPolicyEp.Req.IEGMISFRPFAEFDEIERITOOD
Search RoomGPT-6 Astra6315059150613800
Claude Fable 5.1631606816065400
Locked StorageGPT-6 Astra4345242940222000
Claude Fable 5.1555824402092600
Blackout SearchGPT-6 Astra551315015011131100
Claude Fable 5.1551296517057600
Painted CubesGPT-6 Astra378558150410220
Claude Fable 5.1471326170012300
Marked MugsGPT-6 Astra491282007024221
Claude Fable 5.1541874104213491
Unfamiliar ContainersGPT-6 Astra511186601416463
Claude Fable 5.1541387201413460
Stamp CompositionGPT-6 Astra4212352703100370
Claude Fable 5.151170151102400490
Odd ParcelGPT-6 Astra46754105900000
Claude Fable 5.146946004000000
Puzzle BoxGPT-6 Astra220001000000
Claude Fable 5.115150053470000
Wobbly StandGPT-6 Astra44440911750050
Claude Fable 5.1464602116700200

Failure breakdown by category

#
Failure breakdown by taxonomy category, in % of the unmet requirements of failed episodes. Overall: 95% bootstrap intervals. the attribution protocol gives the attribution rules.
SearchInspectTestOverall
PhaseCategoryLinked capabilitiesAstraFableAstraFableAstraFableAstraFable
Before actionInadequate evidence gatheringAcquisition, inference50.659.739.040.915.225.236.7 · [34, 40]42.2 · [39, 45]
Misinterpretation of evidenceUse16.620.53.92.415.26.211.5 · [9, 14]9.0 · [7, 11]
Faulty reasoning or planningUse, inference, organization0.60.34.84.420.115.77.4 · [6, 9]6.4 · [5, 8]
During actionFailed action executionOrganization7.44.628.117.329.924.321.1 · [19, 23]15.4 · [13, 17]
After feedbackFailure to detect errorsUse, organization14.16.33.31.30.00.06.3 · [5, 8]2.5 · [2, 3]
Ineffective error recoveryInference, organization10.78.63.03.10.00.05.0 · [4, 6]3.9 · [3, 5]
Incorrect terminationUse, organization0.00.016.330.419.728.611.3 · [8, 15]20.5 · [17, 24]
Failed episodes161173137155134158432486
Unmet requirements (denominator)3263473314572443259011129

Workload and occlusion contrasts

#
Workload and occlusion on the four tasks that vary both, pooled over both models. Brackets: 95% bootstrap intervals. Last block: workload drop minus occlusion drop.
SRProg
AstraFablePooledAstraFablePooled
Workload
Low21.77.214.453.031.342.1
Mid12.84.48.646.833.039.9
High5.80.02.941.427.134.2
High - Low-15.8 · [-25.6, -6.7]-7.2 · [-11.7, -2.8]-11.5 · [-16.8, -6.2]-11.6 · [-20.5, -2.8]-4.2 · [-11.7, 3.3]-7.9 · [-14.0, -1.8]
Occlusion
Visible16.75.611.151.134.242.7
Look10.84.77.845.627.636.6
Uncover12.81.47.144.529.537.0
Look - Visible-5.8 · [-15.0, 3.3]-0.8 · [-5.6, 5.3]-3.3 · [-8.8, 1.9]-5.6 · [-14.3, 3.5]-6.6 · [-13.8, 0.7]-6.1 · [-11.6, -0.4]
Uncover - Visible-3.9 · [-13.9, 5.6]-4.2 · [-8.3, 0.0]-4.0 · [-9.3, 1.2]-6.6 · [-15.2, 1.8]-4.7 · [-11.9, 2.7]-5.6 · [-11.6, 0.4]
Workload drop minus occlusion drop
vs. Look10.0 · [-3.1, 23.1]6.4 · [1.4, 12.5]8.2 · [1.5, 15.0]6.1 · [-6.0, 18.9]-2.4 · [-12.0, 7.6]1.8 · [-6.3, 9.8]
vs. Uncover11.9 · [-0.6, 24.2]3.1 · [0.0, 7.5]7.5 · [1.1, 13.8]5.0 · [-6.9, 16.9]-0.5 · [-10.1, 9.3]2.3 · [-5.9, 10.5]

Where additional failures occur

#
Where the extra failures come from. Unmet requirements per episode by failure phase, on the tasks and instances of Workload and occlusion contrasts. Brackets: 95% bootstrap intervals.
Before actionDuring actionAfter feedbackAll
Workload
Low0.660.480.491.63
Mid1.010.610.722.36
High1.420.671.103.19
High - Low0.75 · [0.53, 0.96]0.19 · [0.05, 0.33]0.61 · [0.29, 0.95]1.55 · [1.23, 1.89]
Occlusion
Visible0.910.650.642.22
Look1.060.490.902.44
Uncover1.150.610.772.53
Look - Visible0.15 · [-0.08, 0.36]-0.17 · [-0.30, -0.03]0.25 · [-0.05, 0.55]0.22 · [-0.11, 0.61]
Uncover - Visible0.23 · [-0.01, 0.46]-0.04 · [-0.17, 0.09]0.13 · [-0.16, 0.48]0.31 · [-0.07, 0.68]

Puzzle Box with hidden locks

#
Puzzle Box: effect of hiding the locks, on chain lengths 4 and 5. Δ: Covered minus No cover, with 95% bootstrap interval. Last row: Claude Fable 5.1's Δ minus GPT-6 Astra's.
PolicyCoverSRProgDecisions
GPT-6 AstraNo cover94.499.684.7
Covered96.499.5119.5
Δ2.0 · [-8.7, 16.7]-0.1 · [-1.4, 1.0]34.8 · [14.0, 55.8]
Claude Fable 5.1No cover100.0100.0106.7
Covered34.659.1174.6
Δ-65.4 · [-83.7, -46.2]-40.9 · [-56.3, -26.0]67.9 · [46.8, 86.1]
Gemini 3.8 FlashNo cover27.845.2172.2
Covered0.025.5199.8
Δ-27.8 · [-44.4, -11.1]-19.7 · [-37.9, -1.0]27.6 · [8.5, 48.8]
Difference in Δ-67.3 · [-89.8, -44.9]-40.8 · [-56.2, -26.1]33.1 · [5.0, 63.6]

How episodes end

#
How episodes end, in % of 500 episodes per model. Stop, Text answer, and Timeout end without a SUBMIT and all fail. Precision: share of submissions that were correct.
SubmittedNo submit
ModelCorrectWrongStoppedText answerTimeoutSubmit ratePrecision
GPT-6 Astra19.465.65.40.09.685.022.8
Claude Fable 5.110.055.028.65.41.065.015.4
Gemini 3.8 Flash2.055.60.80.041.657.63.5

Workload by task

#
Workload by task. Each cell gives Low / Mid / High. SR, Prog, and T/O (share of episodes that ran out of decisions) are in %; Dec is the mean number of decisions per episode. Workload levels are listed in the methods. Each level has about 17 episodes per model, so single-task trends are noisy.
TaskPolicySR L / M / HProg L / M / HDec L / M / HT/O L / M / H
Stamp CompositionGPT-6 Astra32 / 6 / 667 / 56 / 41104 / 113 / 980 / 0 / 0
Claude Fable 5.10 / 6 / 030 / 45 / 38138 / 142 / 1460 / 0 / 0
Odd ParcelGPT-6 Astra0 / 0 / 017 / 21 / 11171 / 186 / 19221 / 7 / 18
Claude Fable 5.10 / 0 / 04 / 8 / 7179 / 180 / 1900 / 0 / 11
Marked MugsGPT-6 Astra6 / 12 / 1246 / 38 / 44111 / 132 / 1360 / 6 / 6
Claude Fable 5.10 / 0 / 028 / 17 / 5136 / 150 / 1390 / 0 / 0
Painted CubesGPT-6 Astra49 / 33 / 681 / 72 / 6995 / 129 / 1420 / 0 / 0
Claude Fable 5.129 / 12 / 063 / 61 / 59121 / 133 / 1580 / 0 / 0
Unfamiliar ContainersGPT-6 Astra15 / 0 / 041 / 28 / 23183 / 189 / 15110 / 41 / 21
Claude Fable 5.10 / 0 / 017 / 20 / 13170 / 181 / 1500 / 0 / 0
Locked StorageGPT-6 Astra38 / 24 / 2354 / 59 / 45140 / 162 / 1496 / 11 / 6
Claude Fable 5.125 / 0 / 043 / 10 / 13113 / 135 / 1190 / 0 / 0
Task-balanced meanGPT-6 Astra23.1 / 12.6 / 7.751.1 / 45.7 / 39.0134 / 152 / 1456.2 / 10.7 / 8.4
Claude Fable 5.19.0 / 3.0 / 0.030.9 / 27.0 / 22.4143 / 153 / 1500.0 / 0.0 / 1.9

Progress through Puzzle Box

#
Puzzle Box: how far episodes get. Share of episodes (%) that reach each stage, on chain lengths 4 and 5, weighted as in Puzzle Box with hidden locks. All bolts: every chain bolt was released at some point. Extracted: the item left the box. Gave up: the episode ended with a declared Stop or a text answer. Timeout: the episode ran out of decisions.
PolicyCoverAll boltsExtractedSolvedGave upTimeout
GPT-6 AstraNo cover1001009400
Covered100969600
Claude Fable 5.1No cover10010010000
Covered474035650
Gemini 3.8 FlashNo cover393928067
Covered1300095

Painted Cubes: direction of errors

#
Painted Cubes: direction of wrong-bin commitments. Cube counts by final outcome. A positive cube satisfies the rule and belongs in the blue bin. FN: positive in yellow; FP: negative in blue. FNR = FN/(TP+FN) and FPR = FP/(TN+FP) over committed cubes, in %, with 95% bootstrap intervals over episodes. Uncaptured cubes (positive / negative) are counted separately.
PolicyRuleTPFNFNRTNFPFPRUncaptured
GPT-6 AstraHas a colour2114.5 · [0.0, 16.7]4947.5 · [0.0, 17.0]7 / 4
Has a shape10533.3 · [12.5, 55.6]3039.1 · [2.5, 19.4]2 / 2
Has a colour + shape91155.0 · [23.8, 82.6]3637.7 · [0.0, 17.1]2 / 7
Exactly one mark2513.8 · [0.0, 12.0]42714.3 · [4.3, 25.0]7 / 14
All rules651821.7 · [11.9, 31.7]157179.8 · [5.5, 14.7]18 / 27
Claude Fable 5.1Has a colour11945.0 · [23.1, 64.3]5012.0 · [0.0, 5.8]9 / 6
Has a shape6857.1 · [20.0, 83.3]2600.0 · [0.0, 0.0]3 / 9
Has a colour + shape91052.6 · [27.3, 78.6]30514.3 · [2.7, 30.6]3 / 11
Exactly one mark15834.8 · [21.1, 48.0]3749.8 · [2.4, 18.9]10 / 22
All rules413546.1 · [34.2, 57.3]143106.5 · [2.9, 10.7]25 / 48
Download all table data ↓