Lab notes
No notes match your search.
2026
ScrambleToolBench: the next step is already in the map
An agent discovers what its tools do. Their names change, and it starts searching again—even when its earlier observations point to the next call.
Read note
MNIST-PRO: seeing the strokes is not enough
When an agent gets a digit wrong, did it miss the evidence or struggle to put the glimpses together? We use MNIST-PRO to tell these failures apart.
Read note
IDEAgent: Agentic Quality-Diversity Search
IDEAgent searches for a diverse set of candidate research ideas and records how each candidate changes during refinement.
Read note
Toward Efficient Data-Centric Training, Part II
PODS changes how much data is used at each stage of training instead of fixing one selection ratio.
Read note
Toward Efficient Data-Centric Training, Part I
Data Agent updates its data-selection policy as the target model trains.
Read note
δ-mem: Efficient Online Memory for LLMs
δ-mem keeps a compact state during inference and uses it to modify attention in a frozen language model.
Read note