RoboQuestDeCLaRe Lab

METHODS & SCORING

Inside the benchmark.

The observation and control contract, learned-policy setup, task-specific scoring and failure analysis.

System Prompt and Tools

Every model receives the same system prompt, sent once at the start of the episode. Only the goal differs between tasks; the goal text is identical across a task’s occlusion variants. The prompt for Puzzle Box reads:

You control a simulated mobile manipulation robot. Use only the provided camera images, robot proprioception and action/budget feedback. Determine what the goal requires from the scene. Choose exactly one arm, base, wait or stop tool per response. There are no semantic object-location or grasp tools. Images are upright RGB from three fixed cameras; new views require physical robot or object motion. The arm tool controls the right gripper site in world metres, with optional world orientation as quaternion xyzw. World z is up. Gripper -1 opens, +1 closes and 0 keeps its current setpoint. The base tool takes normalized body-frame forward, left and yaw velocity commands; positive yaw turns left. The arm stays relative to the base during base motion. The torso is held. A tick is 0.05 seconds. Commands consume up to 200 ticks each, ending early on the first physical SUBMIT button press. Check observed outcomes: commanded targets are not guaranteed to be reached. Physically press SUBMIT to commit the result; the first press freezes the score and ends the episode immediately. Wrong submission, timeout without submission, or calling stop without submission fails. The stop tool abandons the episode without advancing time. Invalid or multiple tool calls perform no physics and still consume this model response. No intermediate success is awarded. Goal: Open the wooden puzzle box on the counter, take out the tangerine inside, and stand it inside the tray. The lid slides sideways but is locked by sliding bolts with round black handles: a part moves only when nothing blocks it, so work out the order. Any order that works is fine. Once the tangerine is inside the tray and you have let go of everything, press the red SUBMIT button. Your first submission ends the episode; a wrong submission or no submission fails. Control contract: mobile-native-control-v1.

Action tools lists the four tools. After each call the model receives a short result with three fields: whether the command was accepted, the current tick count, and an error that is either empty or the fixed message “Invalid action; check schema and remaining budget.” The result never contains simulator exceptions, scores, or task state.

View the full table.

Evaluation Harness

Lineage.

The harness builds on the Inspect Robots agent plugin (plugin version 0.26.0, MIT license), the framework used for the public GPT-6 Astra real-robot runs. We reuse its provider clients for the OpenAI Responses and Anthropic Messages interfaces, their message formats, native reasoning continuation (encrypted reasoning items and thinking signatures), its rule for dropping old images, and its pattern of one tool call per turn. The simulator (RoboCasa kitchens in MuJoCo, with a Panda arm on a mobile base), the observation filter, the action tools and their execution, the submission protocol, the system prompt, the budget rules, retry and restart handling, prompt caching, a native Gemini client, and all logging and scoring are ours.

Observations.

Each decision adds one user message. Its text part gives the end-effector and base poses, the 13 joint positions and velocities rounded to millimetres and milliradians, the elapsed ticks and simulated seconds, and the number of model responses left, counting down from 200. Its image part holds three labelled 512×512512\times512 RGB images from the left and right scene cameras and the wrist camera, without overlays. A positive allowlist builds every observation, so no task specification, seed, instance identifier, or hidden ground truth can reach the model.

Actions and execution.

The model must reply with exactly one tool call (Action tools) and chooses its duration in ticks. An arm command runs a proportional servo on the measured gripper pose through the robot’s operational-space controller for the requested number of ticks, a base command applies the velocity each tick with the arm held relative to the base, and a wait holds both. A command runs for exactly its requested ticks, with no added settling time, and stops early only when the SUBMIT button is pressed. A reply with no tool call, several tool calls, or arguments outside the schema runs no physics but still uses up the decision; a reply without any tool call also ends the episode.

Memory.

Every episode starts a fresh conversation, with no memory across episodes and no notes from earlier runs. Within an episode the conversation only grows: the system prompt, then for every decision the observation, the model’s full reply including its native reasoning content, and the tool result. Text is kept for all decisions. Images are removed from all but the two most recent observations, so six images are in context at a time, while the text of older observations, and with it the full proprioceptive history, stays. Nothing is summarized. To keep costs manageable, the Anthropic and OpenAI clients mark cache breakpoints on the newest observations whose images have been dropped, so each turn reads the prefix the previous turn wrote; the Gemini client replays the full history with every turn’s thought signature. If a provider rejects a reasoning continuation, the harness archives the conversation and restarts it from the system prompt and the current observation, at most twice per episode and without replaying any physics.

Budgets and termination.

An episode allows 200 model decisions and at most 200 ticks per command. Each request may be retried up to six times, and retries do not count as decisions. The tasks also define a budget in ticks, but it is neither shown to the model nor enforced in these runs, so the decision limit is the binding one. An episode ends at the first SUBMIT press, after 200 decisions, or when the model calls stop or replies without a tool call. Only the state at the press is scored, and the scene freezes that score when the button is pressed. Episodes interrupted by the provider or the runtime are left unscored rather than counted as failures, and each reply’s model identifier is checked before its action runs.

Changes from Inspect Robots.

The published GPT-6 Astra configuration of Inspect Robots allows 20 model calls with images on every turn and computes motion durations from speed limits. We change it in the following ways.

  1. Physical submission replaces done and give_up. Upstream scores success when the model declares it or when time runs out, so a policy can reach a correct state by trying several answers. Pressing a physical button commits a single answer, which is what makes information gathering meaningful: the policy has to decide when it knows enough.

  2. The model chooses durations, and nothing settles automatically. Every tick is charged, and elapsed time is visible in each observation, so managing time is part of the policy’s task.

  3. A new tool set. Upstream’s generated tools only support its own rotation format. We use absolute world poses with quaternions, add base and wait tools, and remove the mandatory note, the picture request, and the done and give_up tools. There are no semantic helpers, so every view has to be earned by moving.

  4. The goal is stated once. The goal and the control contract are in the system prompt rather than repeated in every observation, and proprioception is rounded, which keeps each turn short and the prompt cacheable.

  5. The decision budget is visible. The model sees how many responses remain but no countdown of physical time, since a countdown would change when a policy chooses to submit.

  6. Observations only. Upstream supports operator feedback, hindsight notes, and learnings carried between trials. We leave all of them out, so the benchmark measures what a policy does within one episode.

  7. Accounting and robustness. One accepted response counts as one decision, retries are counted separately, truncated or invalid replies never execute a partial action, continuation failures are recovered by restarting the conversation, and token usage and cost are recorded for every provider.

  8. Prompt caching. With 200 decisions of three images each, re-sending the conversation dominates the cost, so cache breakpoints let each turn reuse the previous prefix.

  9. Scale. Episodes allow 200 decisions instead of 20, with 512×512512\times512 images from three cameras, including the wrist.

Every episode in the release keeps its system prompt, tool schemas, instruction, configuration, full and archived conversations, per-decision images, a tick-level physics trace, a video, and its token usage and cost.

π0.5\pi_{0.5} Fine-Tuning and Evaluation Details

In addition to evaluating zero-shot frontier language-vision models (the benchmark description), we evaluate a supervised fine-tuned π0.5\pi_{0.5} model. This section summarises the observation mapping, action representations, training objective, and hyperparameters used.

Observation Mapping.

At each decision step, the model ingests:

Action Space and Horizon.

The action expert generates an action chunk of length H=20H=20 time steps, which corresponds to 1.01.0 second of continuous execution at the simulator’s native 2020 Hz control frequency. Each action predicts a 12-dimensional hybrid mobile manipulation command: 6-D end-effector delta pose, 1-D continuous gripper action (±1\pm 1), 3-D mobile base velocities (forward, lateral, and yaw), and 2-D torso pose.

Training Objective and Joint Modeling.

In the joint formulation, training simultaneously optimizes an autoregressive language objective for high-level subtask prediction and a conditional flow-matching objective for low-level action chunking: ℒtotal=ℒsubtask+10⋅ℒflow,\mathcal{L}_{\text{total}} = \mathcal{L}_{\text{subtask}} + 10 \cdot \mathcal{L}_{\text{flow}}, where ℒsubtask\mathcal{L}_{\text{subtask}} is standard cross-entropy loss over causal subtask tokens, and ℒflow\mathcal{L}_{\text{flow}} is the flow-matching loss conditioned on vision, state, prompt, and the teacher-forced subtask.

Hyperparameters.

Training and inference settings summarizes the key training and inference hyperparameters. Fine-tuning uses the AdamW optimizer with β1=0.9\beta_1=0.9, β2=0.999\beta_2=0.999, weight decay 1×10−41\times 10^{-4}, and gradient norm clipping at 1.01.0. The learning rate follows a cosine decay schedule with a linear warmup of 1000 steps, starting from a peak learning rate of 2.5×10−52.5\times 10^{-5} down to a minimum of 2.5×10−62.5\times 10^{-6}. Global batch size is set to 32 using full-parameter training with exponential moving average (EMA) weight tracking enabled.

View the full table.

During evaluation, the policy runs in closed loop, replanning every 20 ticks (1.01.0,s) until the episode terminates upon pressing the physical SUBMIT button or exceeding the task horizon.

Failure Taxonomy Details and Link to Capabilities

Inadequate evidence gathering concerns missing information where the agent prematurely acts decisively before obtaining key evidence for a reliable decision. For example, it may leave a relevant location unexplored, fail to unveil a distinguishing information, or omit a diagnostic test. In Marked Mugs, placing a vessel before inspecting the label underneath would be an instance of such failure. Lacking in evidence acquisition and/or interactive inference capabilities curtail agent’s capability to gather enough key evidences.

Misinterpretation of evidence acquired in the context of the task may lead to erroneous downstream decisions. For instance, the agent observes the correct object, yet misidentifies it, misreads a mark on it, or misunderstands the result of an interaction with it. In Marked Mugs, misreading the label underneath a vessel would likely lead to a wrong placement of it, precipitating into a task failure; a possible manifestation of the lack of evidence usage capabilities. Acquisition and interpretation therefore fail at different points. The former lacks the needed observation and the latter assigns a wrong meaning leading to a task failure.

Faulty reasoning or planning around the gathered evidences may produce an ineffective or erroneous plan. An agent may lose track of a relevant earlier result, use stale information, omit a goal constraint, make incorrect inferences from prior results, or choose an invalid action order. For example, in Odd Parcel task, identifying the odd parcel requires combining the results of balance comparisons before identifying the parcel. On the other hand, correct evidence gathering can still lead to an incorrect plan due to reasoning failure. A lack of evidence use, interactive inference, and/or action organization capabilities may contribute to such failures.

Failed action execution concerns the failure to achieve the intended result from an executed action. For instance, a grasp can miss, an object can fall during transport, or it can be misplaced. This is very likely a fault in action organization capabilities.

Failure to detect errors that matter to task completion may keep the agent away from task progress and completion—likely, a combination of faults in evidence acquisition, use, and action organization. For instance, in Wobbly Stand task, the failure to gauge the tilt of the stand by tracking the ball on top would inexorably lead to task failure.

Ineffective error recovery may also perpetually keep an agent away from progressing toward the goal, that may stem from impaired interactive inference and/or action organization. For instance, a fallen object of interest that remains on the floor despite repeated retrieval attempts may stop meaningful progression. Recovery requires both detecting the discrepancy and thereafter acting toward a corrected state.

Incorrect termination involves submitting at an end state misjudged as success or prematurely abandoning a goal that remains reachable, possibly a case of a lack of evidence use or action organization. A dropped tomato followed by unsuccessful retrieval and pressing SUBMIT illustrates a chain across execution, recovery, and completion. Reaching the SUBMIT button completes that action while leaving the original goal unfinished.

Fine-Grained Capabilities for RoboQuest

Directed search is the core of the Search family where the agent chooses where to look, track where it has checked, and therefrom determine next likely search zone. An unobserved item is not necessarily absent, so an empty location counts only if the agent has actually checked it. For instance, in Search Room, the agent must judge when the remaining search space is exhausted, because the number of storage places and which of them are empty are not given in the instruction. Locked Storage adds dependency: finding a token changes which places can be searched next, so the search order must follow a chain the agent discovers along the way. Blackout Search removes free observation: the agent sees only what its light source illuminates, and it must put the light down to grasp the target; it then acts from memory in the dark. All three tasks require evidence integration, since each observation matters only in combination with what the agent has already seen.

Active perception is the core of the Inspect family where the agent actively scans for pertinent information through interaction, informing next pivotal task actions. In Painted Cubes, the target may be a cube with exactly one painted face. Two painted faces rule a cube out. One painted face does not confirm it while other faces remain hidden. A partial view can therefore reject a candidate, but it cannot confirm one, so the agent memorize the prior observations of the faces. Memory is thus a key capability to keep track of the progress and retrieve relevant information from past observations and inferences. In Marked Mugs, the label under each container can be read only by lifting or tilting it and decide next action. In Unfamiliar Containers, the hidden property is the opening mechanism itself. The agent must discover the affordances of each container through interaction, and it must recognize that some containers cannot be opened at all.

Experimental identification is the core of the Test family where the agent chooses trials that separate candidate explanations and adapt latter trials to earlier outcomes. In Stamp Composition, trial impressions on the test surface reveal stamp patterns and rotations. In Odd Parcel, a two-pan balance is the only instrument. One-by-one weighing is valid but slow, while group weighing identify the odd parcel in much fewer comparisons. The task therefore measures the quality of experimental design, not only its outcome. Physical causal inference links an action to its physical consequences. In Puzzle Box, the agent learns the release order by trying moves and observing which parts move. Puzzle Box is the simplest case of this capability, and the strongest model solves most of its episodes. In Wobbly Stand, the tilt is too small to see. A ball placed on the surface rolls toward the low side and reveals it, and the agent must then relate each shim to the change it produces. These tasks test local inference from physical consequences, not general causal understanding.

Gathering information also has physical consequences. In Marked Mugs, reading a label may spill the contents, and the agent must recollect them before it submits. In Locked Storage, closing a compartment with a token inside locks that token away for good. In Puzzle Box, some wrong moves jam another part until they are reversed. Stamp Composition offers the opposite case. It provides a test surface where trials are cheap, while marks on the final surface are permanent. Taking such consequential action requires the agent to weigh the value of more information against its physical cost, and to judge when the evidence is strong enough to act. Every task ends with an explicit SUBMIT, so this judgment is part of every episode.

Long-horizon planning is an often essential in agentic setting. In Locked Storage task, for example, a sequence of actions need to thought in advance: finding a token →\rightarrow obtaining it without →\rightarrow placing at target location →\rightarrow opening a new storage →\rightarrow obtain new token without closing storage →…\rightarrow \dots.

Task progress

Every episode ends at the first physical press of the SUBMIT button, or when the policy stops or exhausts its budget; the scene state at that moment is frozen and scored. Two quantities are reported per episode: success (the task predicate holds at the press) and progress P∈[0,1]P\in[0,1]. PP measures the state of the task’s own objects, staged as the success predicate is staged; means such as opening a lid or carrying the lantern are never credited. A stage counts only in a passive state: the object released, supported and, where the predicate requires it, untouched by the robot. PP is not capped, so P=1P=1 states that the goal configuration is reached; success additionally requires the task’s tidiness clause, where one exists, and a valid press. One tidiness clause, the deduction δ=0.1\delta=0.1, appears in three tasks; every other side condition is recorded as a flag, not as a term. PP is clipped to [0,1][0,1] after the deduction.

Two rules are shared by several tasks. Three-stage object rule (search room, locked storage, blackout search, unfamiliar containers): for each object ii of nn, the stage si∈{0,1,2,3}s_i\in\{0,1,2,3\} is 11 once its hiding place has been opened, 22 once the object has been out of its hiding place, 33 once it is placed at its destination, and P=1n∑i=1nsi3.P=\frac{1}{n}\sum_{i=1}^{n}\frac{s_i}{3}. \label{eq:three-stage} Stages 1 and 2 are sticky: an object that has been out keeps si≥2s_i\ge 2 even if a closer shuts it back in or it falls to the floor. A cloche target reaches si=2s_i=2 as soon as the cloche is lifted; an open-counter target has no stage 1. Chain rule (puzzle box, locked storage): each lock or bolt of the chain is one share and the object’s stages together are one share; once the object is out, the chain counts as fully released.

Search room.

Equation [eq:three-stage] over the nn targets, with stage 1 = the target’s compartment opened to at least 15%15\,\% of its joint range (drawers 0.150.15 m, doors 0.240.24 rad), stage 2 = out of its hiding place (out of the compartment interior, out from under the cloche, or moved 0.150.15 m from its open spot), stage 3 = on the tray, supported and upright, less δ\delta if a non-target object overlaps the tray at the press. Flags: target dropped to the floor, closer re-shut a target, objects not still.

Locked storage.

Let dd be the chain depth, uu the number of chain links currently passable (u=du=d once the target has been out; the dead-end compartment never counts), and ss the target’s stage as in search room. Then P=max(0,u+s/3d+1−δ[a non-target object overlaps the tray]).P=\max\left(0,\;\frac{u+s/3}{d+1}-\delta\,[\text{a non-target object overlaps the tray}]\right). Flags: relocks, token on the floor, dead end opened.

Blackout search.

As search room, including the deduction; the lantern contributes nothing. Flags: lantern lifted, carried, on the floor; lantern distance at each opening.

Marked mugs.

For each of the nn vessels let vi=1v_i=1 if it stands upright on the pad of its label colour with its own two balls inside, vi=0.5v_i=0.5 if it stands upright on that pad but the balls are not both its own, and vi=0v_i=0 otherwise (a wrong pad is 00, not a penalty). Then P=1n∑iviP=\frac1n\sum_i v_i. Flags: vessels on a wrong pad; other objects released and still.

Painted cubes.

With nn cubes, cc captured in their correct bin and ww captured in the wrong bin (captures are irreversible), P=cn.P=\frac{c}{n}. Flags: wrong captures ww, coverage (c+w)/n(c+w)/n.

Unfamiliar containers.

Equation [eq:three-stage] over the nn items, with stage 1 = the box holding the item opened once, stage 2 = the item out of its box, stage 3 = the item in the bowl. Flags: items released; the box mechanism type, reported against its opening rate.

Stamp composition.

Let CC be the set of target cells of the reference pattern and, for c∈Cc\in C, let gc=1g_c=1 if the cell holds only clean dots (each centre within 99 mm of the cell centre) and no more than the design’s kck_c dots, gc=0.5g_c=0.5 if a dot crosses the cell edge or the cell holds more than kck_c dots, and gc=0g_c=0 if the cell is empty. Let mm be the number of stray dots, i.e. dots on the paper outside CC, plus one if the ink outside the grid exceeds 3535 pixels. Then P=∑c∈Cgc|C|+m.P=\frac{\sum_{c\in C} g_c}{|C|+m}. Dots are located from the recorded impressions with the stamp’s die geometry. Flag: every stamp released, still and supported (required for success). An untouched scene reads P=0P=0.

Odd parcel.

Let OO be the set of odd parcels and BB the set of parcels released inside the answer box. Then P=max(0,|B∩O||B∪O|−δ[a parcel rests on either pan, or the balance was moved]).P=\max\left(0,\;\frac{|B\cap O|}{|B\cup O|}-\delta\,[\text{a parcel rests on either pan, or the balance was moved}]\right). An empty box reads 00; the exact odd set with parcels left on the pans reads 0.90.9.

Puzzle box.

Let LL be the chain length, rr the number of chain bolts currently released (a relocked bolt does not count while the item is inside; r=Lr=L once the item has been out; decoy sliders never count), and s∈{0,1,2,3}s\in\{0,1,2,3\} the item stage: 11 once the lid has slid past half its travel, 22 once the item has been out of the cavity, 33 once it stands in the tray, supported and released. Then P=r+s/3L+1.P=\frac{r+s/3}{L+1}. Flags: robot contact with any part at the press (required absent for success), relock events, the lock at which a failed episode stalled.

Wobbly stand.

Let θ0\theta_0 be the initial tilt of the stand and θ\theta its tilt at the press. Let q=clip(1−θ/θ0,0,1)q=\operatorname{clip}(1-\theta/\theta_0,\,0,\,1) while the stand is released, settled and untouched, q=1q=1 once θ\theta is within the level threshold, and q=0q=0 while the robot holds or touches the stand (values below the deadband 0.010.01 read 00). Let b=1b=1 if the ball is on top of the stand, else 00. Then P=q1+b2.P=q\,\frac{1+b}{2}. A level stand without its ball reads 0.50.5; an untouched scene reads 00. Flags: ball fell off, ball put back.

View the full table.

Failure Attribution Details

Cohort.

The attribution uses every instance that both GPT-6 Astra and Claude Fable 5.1 played with an official score, after removing seed-zero runs, superseded registry entries, runs flagged as scene suspects, and runs interrupted by the provider or the runtime. This gives 540 instances per model, between 46 and 63 per task. It differs slightly from the 50 instances per task of Success and task progress, so rates here are not directly comparable to that table. For Stamp Composition we use the dot rule of the benchmark description for success. Five GPT-6 Astra Odd Parcel episodes and one Claude Fable 5.1 episode boxed exactly the odd set but left a parcel on a pan. The instruction does not ask for empty pans, so we report these separately and leave them out of the breakdown.

Signals from the action log.

Every observation reports the gripper pose and the two finger joints. A close command after which the fingers stay at least 6 mm apart is a grasp, and one after which they close to under 4 mm is a miss. A grip is lost when the fingers collapse below 4 mm while the close command still holds, and an open command while holding is a release. A hold is a carry when the gripper lifts by more than 6 cm and moves by more than 15 cm, which separates carrying an object from pulling a handle. A repair attempt is any later grasp or miss within 30 cm of the point where a carried object was lost. We follow objects through carries by matching each grasp to the nearest tracked object within 8 to 11 cm. Thresholds were set on a few episodes and then checked on video frames.

Rules.

Each unmet requirement gets exactly one category. When a unit fails in execution and a response can be observed, it is attributed to the response: no repair attempt gives failure to detect errors, and a failed repair attempt gives ineffective error recovery. A requirement never attempted in an episode that ended by SUBMIT, Stop, or a text answer is incorrect termination. In an episode that ran out of decisions it is out of decisions. Failure attribution rules lists the task-specific rules.

View the full table.

Validation and limits.

The earlier review of the failed episodes’ videos coded, for each episode, whether an observed execution failure caused the error and whether a mark was misread. On the first question our attribution agrees on 195 of 217 episodes where the review reached a verdict (κ=0.79\kappa=0.79). On the second it agrees on 43 of 50 episodes (κ=0.63\kappa=0.63); most of the disagreements are targets that the gripper reached but did not collect, which we count as misinterpretation and the review left open. The rules have known biases. Inadequate evidence gathering in the Search family is an upper bound, since a target seen from afar and never approached also counts. Misinterpretation in Painted Cubes is a lower bound, since a missed face explains some errors that were in fact misreads. The Odd Parcel split between IEG and FRP rests on the number of pan placements, not on which parcels were compared. No rule uses the agent’s own text, because GPT-6 Astra returns no readable reasoning and text would favour the model that writes more. The code and the episode-level attributions are released with the benchmark.

View the full table.

A Detailed Analysis on Exploration Impairment

the benchmark description argues that RoboQuest tasks demand four groups of capabilities (Capabilities required by each task): evidence acquisition, evidence use, interactive inference, and action organization, and that a weakness in any of them may lead to the failures in our taxonomy (the benchmark description). the benchmark description showed that the frontier models fail most episodes. In this section, we connect the two. We first assign every failure of GPT-6 Astra and Claude Fable 5.1 to a taxonomy category, and then read the result one capability group at a time.

Attributing failures.

A failed episode can fail in several ways at once. A Search Room episode with three targets may end with two of them missing, and a Painted Cubes episode with eight cubes may misfile three. We therefore count failures per unmet requirement rather than per episode: every target left out of the tray, every misfiled or unsorted cube, every vessel not correctly placed, every wrong stamp cell, every missing or extra parcel, and, for Puzzle Box and Wobbly Stand, the single unsolved chain or stand (Failure attribution rules). We assign each unmet requirement to the first step of the workflow that failed and was not corrected. The assignment uses the final simulator state, the hidden task specification, and the gripper trajectory recorded at every decision. For example, a target whose hiding place the gripper never approached is an inadequate evidence gathering failure, while a target that was picked up and then dropped is a failed action execution. The rules are the same for both models (the benchmark description), and on whether an execution error caused the failure, they agree with a manual review of the episode videos on 90% of the episodes. Failure breakdown by category summarizes the result by task family, with the capability groups linked to each failure type, and the benchmark description breaks it down per task.

More than half of all failures occur before the agent acts (56% for GPT-6 Astra and 58% for Claude Fable 5.1), about a fifth during action (21% and 15%), and the rest after feedback (23% and 27%). The two models therefore share the same main weakness, but they differ after that. GPT-6 Astra fails more often in execution and in reacting to its own errors, while Claude Fable 5.1 fails more often by stopping early. We now discuss each capability group.

View the full table.

Evidence Acquisition

Directed search and active perception decide whether the agent obtains the evidence it needs. Their failure, inadequate evidence gathering, is the largest single category for both models (37% and 42%). In Search Room and Blackout Search, the gripper never came within 25 cm of the hiding place of 83% (GPT-6 Astra) and 91% (Claude Fable 5.1) of these missing targets. In Unfamiliar Containers, more than half of the missing items stayed in boxes that the robot never approached. Hence, the agents rarely miss a target they have found. Instead, they stop searching while relevant places remain unchecked.

Occlusion.

Occlusion adds a search or uncovering step before the task. In Visible, the relevant objects are in view; in Look, the agent must first find them from another viewpoint; and in Uncover, it must first physically uncover them. Four tasks vary it: Stamp Composition, Odd Parcel, Marked Mugs and Painted Cubes. Compared with Visible, Look and Uncover lower success by only 3.3 and 4.0 points on average over the two models, and neither interval excludes zero, while progress falls by 6.1 and 5.6 points (Workload and occlusion contrasts). Occlusion mainly moves failures into evidence acquisition: Look and Uncover add 0.15 and 0.23 failures before action per episode, and Look removes 0.17 failures during action (Where additional failures occur). An object that is never found is never handled, so its failure appears as missing evidence rather than as a failed action.

Painted Cubes.

Painted Cubes isolates active perception, since the deciding marks can be on any face of a cube. The agent must put each cube that satisfies a rule into the blue bin and every other cube into the yellow bin. A qualifying cube in the yellow bin is a false negative (FN), and a non-qualifying cube in the blue bin is a false positive (FP). GPT-6 Astra makes more false negatives than false positives, with a false-negative rate (FNR) of 21.7% against a false-positive rate (FPR) of 9.8%, and Claude Fable 5.1 does so far more, 46.1% against 6.5% (the benchmark description). The gap between the two rates is 27.6 points larger for Claude Fable 5.1 (95% CI [10.9,43.4][10.9, 43.4]). This direction points to incomplete inspection. When the rule asks for a mark, an agent that skips a face can wrongly reject a qualifying cube, but it can wrongly accept a cube only by misreading a mark. Unseen faces can explain 39 of the 48 wrong bin choices of Claude Fable 5.1, against 22 of 35 for GPT-6 Astra. Claude Fable 5.1 also left 42 hidden cubes unfound and 27 visible cubes untouched, against 27 and 13 for GPT-6 Astra. Thus, Claude Fable 5.1 stops inspecting too early, both on each cube and across the scene.

View the full table.

Evidence Use

Once the evidence is in view, evidence integration and memory decide whether it is used correctly. Misinterpretation of evidence accounts for 11.5% and 9.0% of the failures. It is concentrated in tasks with look-alike objects. In Locked Storage, it accounts for 29% and 40% of the failures, mostly a similar object placed on the tray instead of the target. Faulty reasoning is most common in the Test family (20% and 16%). In Odd Parcel, it accounts for 59% and 40% of the failures: the agents placed parcels on the balance at least as many times as there are parcels, yet boxed a wrong or incomplete set, or nothing. The balance results were available, but they were not combined into a conclusion.

Workload.

Workload is the number of units a task requires: 2, 3 or 4 stamps, 4, 6 or 8 parcels, 3, 4 or 5 vessels, and 4, 6 or 8 cubes (Low, Mid and High). More units mean more evidence to integrate and remember, and a longer plan. From Low to High workload, success falls by 15.8 points for GPT-6 Astra (21.7% to 5.8%) and by 7.2 points for Claude Fable 5.1 (7.2% to 0.0%), with both intervals excluding zero (Workload and occlusion contrasts). The decline is not caused by running out of decisions, since timeouts stay below 6% at every level. Adding Unfamiliar Containers and Locked Storage, which also vary workload, gives the same trend: over all six tasks, success falls from Low to High by 15.4 points for GPT-6 Astra and 9.0 points for Claude Fable 5.1 (Workload by task). The workload drop exceeds the occlusion drops by 8.2 and 7.5 points, as success requires every unit to be right. The taxonomy shows where the extra failures arise (Where additional failures occur). From Low to High, the number of unmet requirements per episode rises by 1.55. Most of the rise occurs before action (0.75) and at termination (0.61), whereas failed execution rises by only 0.19. Hence, the additional units are not lost because they are harder to handle. They are lost because the agent does not inspect, test or attempt them before it stops, which points to limits of evidence integration, memory and long-horizon planning.

View the full table.

Interactive Inference

Affordance discovery, experimental identification and physical causal inference require the agent to learn from the outcome of its own actions. In Unfamiliar Containers, 14% of the failures of both models are boxes that the robot grasped three or more times without opening them, repeating a move that brought no progress. In Wobbly Stand, on the other hand, the ball test mostly works: a shim under the wrong leg, which would indicate a misread roll, occurs only 4 times and once.

Puzzle Box.

Puzzle Box isolates physical causal inference. A chain of sliding bolts locks the lid, and each bolt moves only after the part blocking it is out of the way. A partial or full opaque cover hides some or all bolt engagements, so the agent must infer the blocking order from how the parts move. Covered and uncovered episodes are matched in chain length. Without a cover, GPT-6 Astra and Claude Fable 5.1 are equally strong, solving 94.4% and 100% of the episodes (Puzzle Box with hidden locks). Under the cover, GPT-6 Astra still solves 96.4%, although its episodes grow from 85 to 120 decisions. Claude Fable 5.1, however, drops to 34.6%, and its episodes grow from 107 to 175 decisions. The two cover effects differ by 67.3 points (95% CI [44.9,89.8][44.9, 89.8]). Gemini 3.8 Flash solves 27.8% without a cover and none with it, running out of decisions in 95% of the covered episodes. None of the 15 failures of Claude Fable 5.1 is inadequate evidence gathering, since it always reaches the bolt that has to move next. In 8 episodes, it pulls that bolt three or more times without releasing it, which is faulty reasoning. In the other 7, it grasps the right bolt but does not slide it far enough. Thus, Claude Fable 5.1 finds the right part but fails to infer how it moves, while GPT-6 Astra recovers the hidden structure through interaction.

View the full table.

Action Organization

Even with the right evidence and inference, the agent must carry out consequential actions and respond to their outcomes. Failed action execution accounts for 21% of the failures of GPT-6 Astra and 15% of those of Claude Fable 5.1, mostly in Marked Mugs, Wobbly Stand and Stamp Composition. In Marked Mugs, only 2 and 9 vessels ended on a wrong pad, so the labels were mostly read correctly. The failures come later: vessels set down at their pad without resting on it (71 and 44) or tipped over (18 and 32). In Wobbly Stand, the stand was knocked over in 13 and 14 episodes, and the shim ended beside the leg rather than under it in 13 and 14 more. In Stamp Composition, 30 and 35 episodes left ink outside the final grid. In these tasks, the agents gather the right evidence but fail to act on it precisely.

Responding to feedback.

The two models differ most in how they respond to the results of their actions. After a carried object was lost, GPT-6 Astra never went back for it in 57% of the cases (54 of 94), against 42% (27 of 65) for Claude Fable 5.1, so failure to detect errors is more frequent for GPT-6 Astra (6.3% against 2.5%). Claude Fable 5.1 instead fails more often through incorrect termination: 20.5% of its failures, against 11.3% for GPT-6 Astra, a gap of 9.2 points (95% CI [4,14][4, 14]). These are requirements it never attempted before stopping or submitting. For example, it left 81 vessels untouched in Marked Mugs, in line with its high stop rate in How episodes end.

View the full table.

In summary, both models fail mostly at evidence acquisition. Beyond that, GPT-6 Astra acts more and breaks more, and often leaves its own errors uncorrected, whereas Claude Fable 5.1 gathers less evidence and ends more episodes with work left undone.

Reported π₀.₅ evaluation

The manuscript figure reports five tasks. These results are kept separate from the frontier models’ 50-episode cells; the figure does not report per-task episode counts.

Fine-tuned π₀.₅ evaluation, transcribed from the manuscript figure.
TaskSuccess (%)Progress (%)
Marked Mugs0.01.5
Puzzle Box1.942.7
Wobbly Stand0.01.1
Stamp Composition0.00.0
Odd Parcel0.00.0