Embodied AI · Efficiency · 2026
← Embodied AI researchMemBodied
A robot's next action can depend on what is no longer visible. MemBodied remembers actions and their observed outcomes in a fixed-size recurrent state.
When the current view is not enough
A robot asked to return a block to its original position needs information that may have disappeared from view. Two episodes can arrive at the same image while requiring different next actions. MemBodied studies this gap in vision-language-action policies: how to carry earlier evidence into the next decision without keeping an ever-growing history of frames.
Read, act, observe, write
MemBodied combines recurrent associative memory with a fixed anchor of the initial scene. The policy queries the associative state, and a gated readout enters a dedicated memory token in the action expert. A separate pathway lets the current observation query the initial-scene anchor.
After the robot executes an action chunk, the next observation reveals its outcome. A delayed write associates the policy state with both the action and its observed consequence. Learned retention and write gates control a delta-rule update. The memory pathways are trained through the policy’s action objective; the episodic memory footprint stays constant in episode length for a fixed model configuration.
Memory-dependent control
Across five RMBench tasks, with 50 rollouts per task and policy, the project reports 50.0% mean success for MemBodied and 6.4% for stateless π₀. Frame stacking reaches 14.8% and vanilla recurrent memory 16.8%.
The comparison with video-history memory includes both NativeMEM alone (38.4%) and NativeMEM integrated with MemBodied (45.2%). This separates the contribution of the video encoding from the recurrent memory mechanism. On LIBERO-Long, success rises from 85.2% to 90.6%.
Physical robot trials
On three physical robot tasks, each evaluated over 20 trials per policy, MemBodied succeeds in 16 of 60 trials (26.67%), compared with 2 of 60 (3.33%) for the stateless baseline. These tasks test remembering an original position, arrangement or order. The project page presents selected demonstrations alongside the aggregate results.
The gains show the value of retaining interaction history, while the remaining failures leave substantial room for better control. Constant-size memory also does not imply lossless recall: the learned state must preserve the evidence needed for the task.
Explore the full method, results and robot demonstrations. For a complementary question—how an agent should remember which sources to trust—see BaRe-Mem.
Project and demonstrations
2026 · recurrent associative memory for vision-language-action models
MemBodied
Method, benchmark comparisons, ablations and physical robot demonstrations.