- Safety Will task-specific agents stay within scope?OffTopicEval Does safety survive fine-tuning?RESTA
- Trustworthiness When should agents trust one another?Epistemic Context Learning Do RAG citations support each claim?Trust-Align
- Multimodality Can models infer rules from visual puzzles?Puzzle Prodigies Can models reason directly from pixels?MIRAS
- AI for Science Can agents find diverse research ideas?IDEAgent Do LLMs judge scientific novelty reliably?LLM-as-Judge for Novelty
- Efficiency Can online memory remain compact?delta-mem Can agents learn continually from memory?Sigma-Mem
- Embodied AI Can preference training improve VLA policies?NORA 1.5 Which data bottlenecks limit VLAs?10 Open Challenges for VLAs
Selected research
All researchEmbodied AI · 2026
GE-Act 2.0
GE-Act 2.0 learns to predict how a scene will change, then uses that prediction to control a robot. The study tests one checkpoint on 100 real-robot tasks, without task-specific fine-tuning.
An AgiBot–DeCLaRe Lab collaboration, with Renhang Liu as DeCLaRe's lead core contributor and Soujanya Poria contributing as an academic advisor.
Model and experimentsAudio and music
TangoFlux
A basketball bounces, shoes squeak, a referee blows a whistle. Hear how TangoFlux follows the scene.
Compare the samplesMemory and learning · 2026
δ-mem and Σ-Mem
δ-mem writes earlier context into an 8 × 8 state. Σ-Mem tracks which other agents have been reliable.
Read δ-mem + Σ-MemEmbodied AI · 2025
NORA and NORA-1.5
NORA learns to turn images and instructions into robot actions from 970,000 demonstrations. NORA-1.5 adds preference training: a world model predicts the outcomes of candidate actions, which are then ranked alongside their distance from demonstrated actions.
Models and robot trials →2017–2023
Earlier work
Multimodality · 2017
Tensor Fusion Network (TFN)
TFN combines words, facial expressions and vocal cues so their interactions can change the predicted sentiment.
Multimodality · 2018–2019
MELD + DialogueRNN
Understanding an emotion often means following who said what, how they said it, and what happened earlier.
Multimodality · Efficiency · 2023–2026
Tango, Mustango + JAM
Tango generates sound from descriptions. Mustango adds beat and chord controls; TangoFlux reduces sampling time; JAM uses word timings to generate songs.
Recent highlights
All updates ↗ScrambleToolBench: adapting to hidden tool changes
A terminal benchmark for agents that must infer undocumented tool behaviour and recover when the tool-to-function mapping changes.
Research internships for DeCLaRe students
Four students started internships at Tencent, MiniMax and Meta.
δ-mem featured in VentureBeat
The article covers δ-mem's compact online state and its low-rank interface to a frozen language model.
Embodied Foundational Models funded by CNRS@CREATE and NRF
A 2026–2029 project on generalist vision-language-action models for embodied AI.
Lab Notes
All notes ↗
August 2026 · Varun Gumma
IDEAgent: Agentic Quality-Diversity Search
IDEAgent searches for a diverse set of candidate research ideas and records how each candidate changes during refinement.
May 2026 · Suorong Yang
Toward Efficient Data-Centric Training, Part II
PODS changes how much data is used at each stage of training instead of fixing one selection ratio.
Code, models and data
More releases ↗Reference implementations of emotion-recognition-in-conversation models, including DialogueRNN.
Multimodal multi-party emotion recognition in conversation, used in later audio and multimodal evaluation work.
An open latent-diffusion model for text-to-audio, built on an instruction-tuned LLM text encoder rather than a contrastive one.
A 3B-parameter open generalist vision-language-action model trained on 970k real-world robot demonstrations, built for low inference cost.
Online associative memory for LLMs — a 0.12% parameter addition that modulates attention through a low-rank interface while the backbone stays frozen.
Text-to-audio generation with flow matching and CRPO preference optimisation. Generates 30 seconds of 44.1 kHz audio in a few seconds on a single A40.
About DeCLaRe
DeCLaRe, short for Deep Cognition and Language Research, was founded by Soujanya Poria at the Singapore University of Technology and Design in 2019 with Navonil Majumder, Devamanyu Hazarika and Deepanway Ghosal. The lab moved to Nanyang Technological University in 2025.
The robot forms 宣 (xuān), “to declare,” carrying the lab's name directly into its visual identity.




