RoboQuestby DeCLaRe Lab GitHub

Failure analysis

Where does exploration break down?

A failed episode can leave several requirements unmet. The analysis attributes each unmet requirement to the first uncorrected failure in the interaction sequence.

Failure taxonomy

Seven ways the investigation breaks.

Failures are organized by where the connection between evidence, decision and physical action first breaks and remains uncorrected.

Before action

Acquire → Interpret → Plan

Was the necessary evidence gathered, understood and used?

During action

Execute

Did the command achieve the intended physical result?

After feedback

Detect → Recover → Stop

Was the outcome checked, repaired if needed, and completion judged correctly?

Definitions and examples from the failure taxonomy. Capability links follow the manuscript’s category-level mapping.
Phase Failure category What breaks Example Linked capabilities
Before action Inadequate evidence gathering Needed evidence was never acquired. A location remains unchecked, a hidden face unobserved, or a diagnostic test omitted. Evidence acquisition · Interactive inference
Before action Misinterpretation of evidence The evidence was observed but read incorrectly. A label is misread or a trial outcome assigned the wrong meaning. Evidence use
Before action Faulty reasoning or planning Available evidence leads to an invalid decision or plan. Balance comparisons are combined incorrectly, or a dependency is ignored. Evidence use · Interactive inference · Action organization
During action Failed action execution The intended physical outcome is not achieved. A grasp misses, a carried object falls, or a shim lands beside its leg. Action organization
After feedback Failure to detect errors A consequential error is not recognized. A dropped object is left behind as the agent continues. Evidence use · Action organization
After feedback Ineffective error recovery The agent attempts a repair but does not correct the error. Repeated retrieval attempts leave the fallen object on the floor. Interactive inference · Action organization
After feedback Incorrect termination The agent submits or stops with reachable goals unfinished. A vessel is never handled before the robot ends the episode. Evidence use · Action organization

Inadequate evidence gathering

The largest failure category comes before execution.

36.7%GPT-6 Astra
42.2%Claude Fable 5.1

Share of unmet requirements in failed episodes. Denominators: 901 for Astra and 1,129 for Fable; 95% bootstrap intervals: [34, 40] and [39, 45].

Share of unmet requirements

The chart needs JavaScript. Exact values are in the failure breakdown table below.

Failed episodes in the matched analysis cohort.Percent · 0–50
01 · Evidence acquisition

Search ends too soon.

In Search Room and Blackout Search, the gripper never came within 25 cm of the hiding place of 55% of Astra’s missed targets and 75% of Fable’s. Relevant places remain unchecked.

02 · Evidence use

More items, less success.

From low to high workload, success falls by 15.8 percentage points for Astra and 7.2 for Fable on the four tasks that vary both workload and occlusion.

03 · Interactive inference

Hidden locks separate the agents.

On the covered Puzzle Box subset, Astra reaches 96.4% success. Fable drops from 100% without the cover to 34.6% with it.

04 · Action organization

Evidence still has to become action.

Failed execution accounts for 21.1% of Astra’s unmet requirements and 15.4% of Fable’s. Incorrect termination contributes 11.3% and 20.5%.

A separate analysis cohort. 540 matched instances per model, with 46–63 per task. This differs from the 500-episode main evaluation. The breakdown contains 432 failed episodes for Astra and 486 for Fable; six ambiguous Odd Parcel outcomes are excluded. These percentages are descriptive attributions, not isolated causal effects.

Failure breakdown by category

Open table ↗
Failure breakdown by taxonomy category, in % of the unmet requirements of failed episodes. Overall: 95% bootstrap intervals. The attribution protocol gives the attribution rules.
Search Inspect Test Overall
Phase Category Linked capabilities Astra Fable Astra Fable Astra Fable Astra Fable
Before action Inadequate evidence gathering Acquisition, inference 50.6 59.7 39.0 40.9 15.2 25.2 36.7 · [34, 40] 42.2 · [39, 45]
Misinterpretation of evidence Use 16.6 20.5 3.9 2.4 15.2 6.2 11.5 · [9, 14] 9.0 · [7, 11]
Faulty reasoning or planning Use, inference, organization 0.6 0.3 4.8 4.4 20.1 15.7 7.4 · [6, 9] 6.4 · [5, 8]
During action Failed action execution Organization 7.4 4.6 28.1 17.3 29.9 24.3 21.1 · [19, 23] 15.4 · [13, 17]
After feedback Failure to detect errors Use, organization 14.1 6.3 3.3 1.3 0.0 0.0 6.3 · [5, 8] 2.5 · [2, 3]
Ineffective error recovery Inference, organization 10.7 8.6 3.0 3.1 0.0 0.0 5.0 · [4, 6] 3.9 · [3, 5]
Incorrect termination Use, organization 0.0 0.0 16.3 30.4 19.7 28.6 11.3 · [8, 15] 20.5 · [17, 24]
Failed episodes 161 173 137 155 134 158 432 486
Unmet requirements (denominator) 326 347 331 457 244 325 901 1129

Task-level diagnosis

The failure mix changes with the task.

Search failures are dominated by missing evidence. Marked Mugs and Wobbly Stand more often break during execution. Inspect each task’s distribution below.

Failure mix by task

The chart needs JavaScript. Exact values are in the complete breakdown below.

Shares of unmet requirements within failed episodes in the matched analysis cohort. Displayed percentages are the reported rounded values. Out of decisions (OOD) means the agent exhausted its 200-decision budget without a valid physical submission. Each accepted model response uses one decision, including invalid tool calls; this is not a token limit.

Complete failure breakdown by task +

Failure categories by task

Open table ↗
Failure breakdown per task. Share (%) of unmet requirements in each category. Ep.: failed episodes; Req.: unmet requirements. IEG: inadequate evidence gathering; MIS: misinterpretation of evidence; FRP: faulty reasoning or planning; FAE: failed action execution; FDE: failure to detect errors; IER: ineffective error recovery; IT: incorrect termination; OOD: out of decisions.
Task Policy Ep. Req. IEG MIS FRP FAE FDE IER IT OOD
Search Room GPT-6 Astra 63 150 59 15 0 6 13 8 0 0
Claude Fable 5.1 63 160 68 16 0 6 5 4 0 0
Locked Storage GPT-6 Astra 43 45 24 29 4 0 22 20 0 0
Claude Fable 5.1 55 58 24 40 2 0 9 26 0 0
Blackout Search GPT-6 Astra 55 131 50 15 0 11 13 11 0 0
Claude Fable 5.1 55 129 65 17 0 5 7 6 0 0
Painted Cubes GPT-6 Astra 37 85 58 15 0 4 1 0 22 0
Claude Fable 5.1 47 132 61 7 0 0 1 2 30 0
Marked Mugs GPT-6 Astra 49 128 2 0 0 70 2 4 22 1
Claude Fable 5.1 54 187 4 1 0 42 1 3 49 1
Unfamiliar Containers GPT-6 Astra 51 118 66 0 14 1 6 4 6 3
Claude Fable 5.1 54 138 72 0 14 1 3 4 6 0
Stamp Composition GPT-6 Astra 42 123 5 27 0 31 0 0 37 0
Claude Fable 5.1 51 170 15 11 0 24 0 0 49 0
Odd Parcel GPT-6 Astra 46 75 41 0 59 0 0 0 0 0
Claude Fable 5.1 46 94 60 0 40 0 0 0 0 0
Puzzle Box GPT-6 Astra 2 2 0 0 0 100 0 0 0 0
Claude Fable 5.1 15 15 0 0 53 47 0 0 0 0
Wobbly Stand GPT-6 Astra 44 44 0 9 11 75 0 0 5 0
Claude Fable 5.1 46 46 0 2 11 67 0 0 20 0

Workload & occlusion

More to keep track of. Less task success.

On the four tasks that vary both factors, the pooled success rate decreases from 14.4% at low workload to 2.9% at high workload. Occlusion effects are smaller and less certain.

Workload

Equal-weight task means pooled over Astra and Fable.

Occlusion

The same four tasks and two models.

Low → high workload: −11.5 percentage points, 95% bootstrap interval [−16.8, −6.2]. Visible → Look: −3.3 [−8.8, 1.9]; Visible → Uncover: −4.0 [−9.3, 1.2]. The occlusion intervals include zero.

Success, progress and confidence intervals +

Workload and occlusion contrasts

Open table ↗
Workload and occlusion on the four tasks that vary both, per model and pooled over the two models. Brackets: 95% bootstrap intervals. Last block: workload drop minus occlusion drop.
SR Prog
Astra Fable Pooled Astra Fable Pooled
Workload
Low 21.7 7.2 14.4 53.0 31.3 42.1
Mid 12.8 4.4 8.6 46.8 33.0 39.9
High 5.8 0.0 2.9 41.4 27.1 34.2
High - Low -15.8 · [-25.6, -6.7] -7.2 · [-11.7, -2.8] -11.5 · [-16.8, -6.2] -11.6 · [-20.5, -2.8] -4.2 · [-11.7, 3.3] -7.9 · [-14.0, -1.8]
Occlusion
Visible 16.7 5.6 11.1 51.1 34.2 42.7
Look 10.8 4.7 7.8 45.6 27.6 36.6
Uncover 12.8 1.4 7.1 44.5 29.5 37.0
Look - Visible -5.8 · [-15.0, 3.3] -0.8 · [-5.6, 5.3] -3.3 · [-8.8, 1.9] -5.6 · [-14.3, 3.5] -6.6 · [-13.8, 0.7] -6.1 · [-11.6, -0.4]
Uncover - Visible -3.9 · [-13.9, 5.6] -4.2 · [-8.3, 0.0] -4.0 · [-9.3, 1.2] -6.6 · [-15.2, 1.8] -4.7 · [-11.9, 2.7] -5.6 · [-11.6, 0.4]
Workload drop minus occlusion drop
vs. Look 10.0 · [-3.1, 23.1] 6.4 · [1.4, 12.5] 8.2 · [1.5, 15.0] 6.1 · [-6.0, 18.9] -2.4 · [-12.0, 7.6] 1.8 · [-6.3, 9.8]
vs. Uncover 11.9 · [-0.6, 24.2] 3.1 · [0.0, 7.5] 7.5 · [1.1, 13.8] 5.0 · [-6.9, 16.9] -0.5 · [-10.1, 9.3] 2.3 · [-5.9, 10.5]

Interactive inference

When the locks disappear, the agents diverge.

Covering the mechanism removes a visual shortcut. Astra maintains high success, while Fable’s success falls sharply. These comparisons use Puzzle Box chain lengths four and five.

Hidden-lock results and uncertainty +

Puzzle Box with hidden locks

Open table ↗
Puzzle Box: effect of hiding the locks, on chain lengths 4 and 5. Δ: Covered minus No cover, with 95% bootstrap interval. Last row: Claude Fable 5.1's Δ minus GPT-6 Astra's.
Policy Cover SR Prog Decisions
GPT-6 Astra No cover 94.4 99.6 84.7
Covered 96.4 99.5 119.5
Δ 2.0 · [-8.7, 16.7] -0.1 · [-1.4, 1.0] 34.8 · [14.0, 55.8]
Claude Fable 5.1 No cover 100.0 100.0 106.7
Covered 34.6 59.1 174.6
Δ -65.4 · [-83.7, -46.2] -40.9 · [-56.3, -26.0] 67.9 · [46.8, 86.1]
Gemini 3.8 Flash No cover 27.8 45.2 172.2
Covered 0.0 25.5 199.8
Δ -27.8 · [-44.4, -11.1] -19.7 · [-37.9, -1.0] 27.6 · [8.5, 48.8]
Difference in Δ -67.3 · [-89.8, -44.9] -40.8 · [-56.2, -26.1] 33.1 · [5.0, 63.6]

Submission & stopping

Reaching a state is not a valid submission.

The robot must physically press SUBMIT. A declared stop, a text answer or running out of decisions ends the episode without completing that requirement.

How episodes end

Open table ↗
How episodes end, in % of 500 episodes per model. Stop, Text answer, and Timeout end without a SUBMIT and all fail. Precision: share of submissions that were correct. Episode outcomes by task gives the breakdown per task.
Submitted No submit
Model Correct Wrong Stopped Text answer Timeout Submit rate Precision
GPT-6 Astra 19.4 65.6 5.4 0.0 9.6 85.0 22.8
Claude Fable 5.1 10.0 55.0 28.6 5.4 1.0 65.0 15.4
Gemini 3.8 Flash 2.0 55.6 0.8 0.0 41.6 57.6 3.5
Episode outcomes for every task +

Episode outcomes by task

Open table ↗
Episode outcomes per task. Percentage of episodes by outcome, with absolute counts shown below in parentheses. Cell shading indicates the percentage within each model–task pair (darker indicates a larger percentage). C: correct submit. W: wrong submit. S: declared stop. A: text answer instead of a submit. T: timeout (out of decisions). Tasks are ordered as in Success and task progress.
GPT-6 Astra Claude Fable 5.1 Gemini 3.8 Flash
Family Task C W S A T C W S A T C W S A T
Search Locked Storage 28%(14) 44%(22) 20%(10) 0%(0) 8%(4) 8%(4) 42%(21) 36%(18) 14%(7) 0%(0) 0%(0) 20%(10) 2%(1) 0%(0) 78%(39)
Search Room 0%(0) 56%(28) 8%(4) 0%(0) 36%(18) 0%(0) 56%(28) 34%(17) 8%(4) 2%(1) 2%(1) 32%(16) 0%(0) 0%(0) 66%(33)
Blackout Search 0%(0) 76%(38) 12%(6) 0%(0) 12%(6) 0%(0) 46%(23) 50%(25) 2%(1) 2%(1) 0%(0) 56%(28) 6%(3) 0%(0) 38%(19)
Inspect Painted Cubes 30%(15) 70%(35) 0%(0) 0%(0) 0%(0) 14%(7) 84%(42) 2%(1) 0%(0) 0%(0) 2%(1) 82%(41) 0%(0) 0%(0) 16%(8)
Marked Mugs 10%(5) 82%(41) 4%(2) 0%(0) 4%(2) 0%(0) 80%(40) 20%(10) 0%(0) 0%(0) 0%(0) 88%(44) 0%(0) 0%(0) 12%(6)
Unfamiliar Containers 6%(3) 66%(33) 6%(3) 0%(0) 22%(11) 0%(0) 24%(12) 68%(34) 8%(4) 0%(0) 0%(0) 36%(18) 0%(0) 0%(0) 64%(32)
Test Puzzle Box 96%(48) 4%(2) 0%(0) 0%(0) 0%(0) 70%(35) 0%(0) 14%(7) 16%(8) 0%(0) 12%(6) 6%(3) 0%(0) 0%(0) 82%(41)
Stamp Composition 14%(7) 84%(42) 2%(1) 0%(0) 0%(0) 2%(1) 74%(37) 22%(11) 2%(1) 0%(0) 0%(0) 54%(27) 0%(0) 0%(0) 46%(23)
Wobbly Stand 10%(5) 86%(43) 2%(1) 0%(0) 2%(1) 6%(3) 62%(31) 30%(15) 0%(0) 2%(1) 2%(1) 96%(48) 0%(0) 0%(0) 2%(1)
Odd Parcel 0%(0) 88%(44) 0%(0) 0%(0) 12%(6) 0%(0) 82%(41) 10%(5) 4%(2) 4%(2) 2%(1) 86%(43) 0%(0) 0%(0) 12%(6)
All tasks 19.4%(97) 65.6%(328) 5.4%(27) 0.0%(0) 9.6%(48) 10.0%(50) 55.0%(275) 28.6%(143) 5.4%(27) 1.0%(5) 2.0%(10) 55.6%(278) 0.8%(4) 0.0%(0) 41.6%(208)