- Safety Will task-specific agents stay within scope?OffTopicEval Does safety survive fine-tuning?RESTA
- Trustworthiness When should agents trust one another?Epistemic Context Learning Do RAG citations support each claim?Trust-Align
- Multimodality Can models infer rules from visual puzzles?Puzzle Prodigies Can models reason directly from pixels?MIRAS
- AI for Science Can agents find diverse research ideas?IDEAgent Do LLMs judge scientific novelty reliably?LLM-as-Judge for Novelty
- Efficiency Can online memory remain compact?delta-mem Can agents learn continually from memory?Sigma-Mem
- Embodied AI Can preference training improve VLA policies?NORA 1.5 Which data bottlenecks limit VLAs?10 Open Challenges for VLAs
Selected research
All researchRobot control · 2026
GE-Act 2.0
GE-Act 2.0 learns to predict how a scene will change, then uses that prediction to control a robot. The study tests one checkpoint on 100 real-robot tasks, without task-specific fine-tuning.
An AgiBot–DeCLaRe Lab collaboration, with Renhang Liu as DeCLaRe's lead core contributor and Soujanya Poria contributing as an academic advisor.
Model and experimentsA model can resist a harmful question yet lose that safeguard after fine-tuning, or accept work outside its assigned role. CoU, RESTA, WalledEval and OffTopicEval test and address these different failures.
Tests, methods and code →Audio and music
TangoFlux
PromptA basketball bounces rhythmically on a court, shoes squeak against the floor, and a referee’s whistle cuts through the air.
Audio unavailable. Open the audio file or choose another sample.
Mustango
Prompt summaryLead guitar over a steady strummed acoustic accompaniment.
100 BPM · Chords: G7 → F7 → C7 → G7
Audio unavailable. Open the audio file or choose another sample.
JAM
Generated songs; titles identify the lyric references, not the original artists’ recordings.
Lyrics excerptHere I go again, running around on my…
Audio unavailable. Open the audio file or choose another sample.
Memory and learning · 2026
δ-mem and Σ-Mem
δ-mem writes earlier context into an 8 × 8 state. Σ-Mem tracks which other agents have been reliable.
Read δ-mem + Σ-MemRobot learning · 2025
NORA and NORA-1.5
NORA learns to turn images and instructions into robot actions from 970,000 demonstrations. NORA-1.5 adds preference training: a world model predicts the outcomes of candidate actions, which are then ranked alongside their distance from demonstrated actions.
Models and robot trials →2016–2021
Earlier work
MultimodalityAffective computing2017–2021
Multimodal representation learning
TFN makes interactions between words, facial expressions and voice explicit. MISA separates what these inputs share from what each contributes on its own; Multimodal-InfoMax trains the combined representation to retain information from all three.
MultimodalityAffective computing2017–2019
Emotion in conversations
A sentence's emotion depends on who says it and what came before. MELD supplies labelled conversations with text, speech and video; DialogueRNN tracks each speaker's changing state, while DialogueGCN models the links between their utterances.
MultimodalityAffective computing2016–2019
Sarcasm detection
Our COLING 2016 work used sentiment, emotion and personality cues to recognise sarcasm in tweets. CASCADE added the discussion and the writer's history. MUStARD brought in voice and facial expression, where the joke can be missed by reading the transcript alone.
Recent highlights
All updates ↗ScrambleToolBench: adapting to hidden tool changes
A terminal benchmark for agents that must infer undocumented tool behaviour and recover when the tool-to-function mapping changes.
Research internships for DeCLaRe students
Four students started internships at Tencent, MiniMax and Meta.
δ-mem featured in VentureBeat
The article covers δ-mem's compact online state and its low-rank interface to a frozen language model.
Embodied Foundational Models funded by CNRS@CREATE and NRF
A 2026–2029 project on generalist vision-language-action models for embodied AI.
Lab notes
All notes ↗
September 2026 · Vernon Toh
ScrambleToolBench: the next step is already in the map
An agent discovers what its tools do. Their names change, and it starts searching again—even when its earlier observations point to the next call.
September 2026 · Vernon Toh
MNIST-PRO: seeing the strokes is not enough
When an agent gets a digit wrong, did it miss the evidence or struggle to put the glimpses together? We use MNIST-PRO to tell these failures apart.
August 2026 · Varun Gumma
IDEAgent: Agentic Quality-Diversity Search
IDEAgent searches for a diverse set of candidate research ideas and records how each candidate changes during refinement.
Code, models and data
More releases ↗Reference implementations of emotion-recognition-in-conversation models, including DialogueRNN.
Multimodal multi-party emotion recognition in conversation, used in later audio and multimodal evaluation work.
An open latent-diffusion model for text-to-audio, built on an instruction-tuned LLM text encoder rather than a contrastive one.
A 3B-parameter open generalist vision-language-action model trained on 970k real-world robot demonstrations, built for low inference cost.
Online associative memory for LLMs — a 0.12% parameter addition that modulates attention through a low-rank interface while the backbone stays frozen.
Text-to-audio generation with flow matching and CRPO preference optimisation. Generates 30 seconds of 44.1 kHz audio in a few seconds on a single A40.
About DeCLaRe
DeCLaRe, short for Deep Cognition and Language Research, was founded by Soujanya Poria at the Singapore University of Technology and Design in 2019 with Navonil Majumder, Devamanyu Hazarika and Deepanway Ghosal. The lab moved to Nanyang Technological University in 2025.
The robot forms 宣 (xuān), “to declare,” carrying the lab's name directly into its visual identity.




