RoboQuest — research overview Synthetic narration. Original supplied simulation demonstrations. Search. Inspect. Test. RoboQuest tests generalist physical agents. A robot must search, inspect, and test to discover what it needs to know, then act on that evidence. Where is the target? In Search Room, the target is not simply waiting in view. The robot must open storage, track where it has looked, and retrieve the complete target set. What is hidden from view? In Painted Cubes, appearances are incomplete. The robot must manipulate each object to reveal hidden faces before deciding where it belongs. What will an action reveal? In Stamp Composition, action is an experiment. Trial impressions reveal patterns and rotations. The robot must use those observations before making permanent marks. What must a generalist agent do? RoboQuest probes nine capabilities in four groups. Evidence acquisition requires directed search and active perception. Evidence use requires integration and memory. Interactive inference requires discovering affordances, identifying unknown properties through experiments, and reasoning about physical causes. Action organization requires handling consequential actions and planning over long horizons. Decide when you know enough. Across ten tasks, a mobile manipulator operates in simulated kitchens with three camera views. Frontier agents receive up to two hundred decisions. A physical submit button commits the result. Partial progress is not success. Progress is not completion. The strongest frontier agent, GPT six Astra, succeeds in nineteen point four percent of episodes. Claude Fable achieves ten percent, and Gemini Flash two percent. Without Puzzle Box, their success rates fall to ten point nine, three point three, and zero point nine percent. Fine-tuning still falls short. The benchmark asks for generalist physical agency, without retraining for every task. Even with task specific fine tuning, the evaluated pi zero point five policy succeeds in only one point nine percent of Puzzle Box episodes, and zero percent on four other tasks. This setup falls short of the capabilities these tasks require. Where does the investigation break? The failure taxonomy distinguishes seven categories. Before action: inadequate evidence gathering, misinterpretation, and faulty reasoning or planning. During action: failed execution. After feedback: failure to detect errors, ineffective recovery, and incorrect termination. Each unmet requirement is attributed to the first failure that remains uncorrected. Missing evidence comes first. The failure analysis traces unmet requirements to evidence gathering, interpretation, planning, execution, and recovery. Inadequate evidence gathering accounts for thirty six point seven percent for Astra, and forty two point two percent for Fable. These are requirement shares, not episode success rates. Demonstrations of investigation. The dataset contains three thousand two hundred and eleven successful scripted demonstrations, totaling four hundred and five point five hours. At twenty hertz, records include three camera views, robot state, actions, and subtask labels. Training and evaluation instances do not overlap. The scripted oracles know the hidden specification, but deliberately perform the investigative steps. Far from reliable physical effectiveness. These evaluated agents remain far from reliable physical effectiveness on RoboQuest. The strongest frontier agent makes forty four point nine percent mean progress, but completes only nineteen point four percent of episodes. Partial progress must become reliable, correctly verified physical outcomes. Act to learn. Then act on it. RoboQuest provides ten tasks and over three thousand successful scripted demonstrations. The challenge is to investigate unfamiliar situations, use the evidence, and recognize when the physical goal has truly been achieved.