Acquire → Interpret → Plan
Was the necessary evidence gathered, understood and used?
Failure analysis
A failed episode can leave several requirements unmet. The analysis attributes each unmet requirement to the first uncorrected failure in the interaction sequence.
Failure taxonomy
Failures are organized by where the connection between evidence, decision and physical action first breaks and remains uncorrected.
Was the necessary evidence gathered, understood and used?
Did the command achieve the intended physical result?
Was the outcome checked, repaired if needed, and completion judged correctly?
| Phase | Failure category | What breaks | Example | Linked capabilities |
|---|---|---|---|---|
| Before action | Inadequate evidence gathering | Needed evidence was never acquired. | A location remains unchecked, a hidden face unobserved, or a diagnostic test omitted. | Evidence acquisition · Interactive inference |
| Before action | Misinterpretation of evidence | The evidence was observed but read incorrectly. | A label is misread or a trial outcome assigned the wrong meaning. | Evidence use |
| Before action | Faulty reasoning or planning | Available evidence leads to an invalid decision or plan. | Balance comparisons are combined incorrectly, or a dependency is ignored. | Evidence use · Interactive inference · Action organization |
| During action | Failed action execution | The intended physical outcome is not achieved. | A grasp misses, a carried object falls, or a shim lands beside its leg. | Action organization |
| After feedback | Failure to detect errors | A consequential error is not recognized. | A dropped object is left behind as the agent continues. | Evidence use · Action organization |
| After feedback | Ineffective error recovery | The agent attempts a repair but does not correct the error. | Repeated retrieval attempts leave the fallen object on the floor. | Interactive inference · Action organization |
| After feedback | Incorrect termination | The agent submits or stops with reachable goals unfinished. | A vessel is never handled before the robot ends the episode. | Evidence use · Action organization |
Inadequate evidence gathering
Share of unmet requirements in failed episodes. Denominators: 901 for Astra and 1,129 for Fable; 95% bootstrap intervals: [34, 40] and [39, 45].
In Search Room and Blackout Search, the gripper never came within 25 cm of the hiding place of 55% of Astra’s missed targets and 75% of Fable’s. Relevant places remain unchecked.
From low to high workload, success falls by 15.8 percentage points for Astra and 7.2 for Fable on the four tasks that vary both workload and occlusion.
On the covered Puzzle Box subset, Astra reaches 96.4% success. Fable drops from 100% without the cover to 34.6% with it.
Failed execution accounts for 21.1% of Astra’s unmet requirements and 15.4% of Fable’s. Incorrect termination contributes 11.3% and 20.5%.
A separate analysis cohort. 540 matched instances per model, with 46–63 per task. This differs from the 500-episode main evaluation. The breakdown contains 432 failed episodes for Astra and 486 for Fable; six ambiguous Odd Parcel outcomes are excluded. These percentages are descriptive attributions, not isolated causal effects.
Task-level diagnosis
Search failures are dominated by missing evidence. Marked Mugs and Wobbly Stand more often break during execution. Inspect each task’s distribution below.
Shares of unmet requirements within failed episodes in the matched analysis cohort. Displayed percentages are the reported rounded values. Out of decisions (OOD) means the agent exhausted its 200-decision budget without a valid physical submission. Each accepted model response uses one decision, including invalid tool calls; this is not a token limit.
Workload & occlusion
On the four tasks that vary both factors, the pooled success rate decreases from 14.4% at low workload to 2.9% at high workload. Occlusion effects are smaller and less certain.
Equal-weight task means pooled over Astra and Fable.
The same four tasks and two models.
Low → high workload: −11.5 percentage points, 95% bootstrap interval [−16.8, −6.2]. Visible → Look: −3.3 [−8.8, 1.9]; Visible → Uncover: −4.0 [−9.3, 1.2]. The occlusion intervals include zero.
Submission & stopping
The robot must physically press SUBMIT. A declared stop, a text answer or running out of decisions ends the episode without completing that requirement.