Deep Cognition and Language ResearchNTU Singapore
We study how models connect language with speech, sound, images, video and robot action.
01
Research themes
- Safety Will task-specific agents stay within scope?OffTopicEval Does safety survive fine-tuning?RESTA
- Trustworthiness When should an agent consult its peers?BaRe-Mem Do RAG citations support each claim?Trust-Align
- Multimodality Can models infer rules from visual puzzles?Puzzle Prodigies Can models reason directly from pixels?MIRAS
- AI for Science Can agents find diverse research ideas?IDEAgent Do LLMs judge scientific novelty reliably?LLM-as-Judge for Novelty
- Efficiency Can online memory remain compact?delta-mem Can robot memory stay fixed in size?MemBodied
- Embodied AI Can robots remember what left the frame?MemBodied Can preference training improve VLA policies?NORA 1.5
01
Safety
02
Trustworthiness
When should an agent consult its peers?
BaRe-Mem
Do RAG citations support each claim?
Trust-Align
03
Multimodality
Can models infer rules from visual puzzles?
Puzzle Prodigies
Can models reason directly from pixels?
MIRAS
04
AI for Science
Can agents find diverse research ideas?
IDEAgent
Do LLMs judge scientific novelty reliably?
LLM-as-Judge for Novelty
05
Efficiency
Can online memory remain compact?
delta-mem
Can robot memory stay fixed in size?
MemBodied
06
Embodied AI
02
Selected recent research
All researchMemBodied: memory for what a robot can no longer see
A robot's next action can depend on what is no longer visible. MemBodied remembers actions and their observed outcomes in a fixed-size recurrent state.
MemBodied reads from associative memory before acting, then writes the action together with its observed outcome. A separate anchor preserves the initial scene.
Across five RMBench tasks it reaches 50.0% mean success, against 6.4% for the stateless baseline.
GE-Act 2.0
GE-Act 2.0 learns to predict how a scene will change, then uses that prediction to control a robot. The study tests one checkpoint on 100 real-robot tasks, without task-specific fine-tuning.
An AgiBot–DeCLaRe Lab collaboration, with Renhang Liu as DeCLaRe's lead core contributor and Soujanya Poria contributing as an academic advisor.
Model and experimentsMultimodalityEfficiency2023–2026
TangoFlux
PromptA basketball bounces rhythmically on a court, shoes squeak against the floor, and a referee’s whistle cuts through the air.
Audio unavailable. Open the audio file or choose another sample.
Mustango
Prompt summaryLead guitar over a steady strummed acoustic accompaniment.
100 BPM · Chords: G7 → F7 → C7 → G7
Audio unavailable. Open the audio file or choose another sample.
JAM
Generated songs; titles identify the lyric references, not the original artists’ recordings.
Lyrics excerptHere I go again, running around on my…
Audio unavailable. Open the audio file or choose another sample.
NORA and NORA-1.5
NORA learns to turn images and instructions into robot actions from 970,000 demonstrations. NORA-1.5 adds preference training: a world model predicts the outcomes of candidate actions, which are then ranked alongside their distance from demonstrated actions.
Models and robot trialsWhen should an agent trust its peers, and when should it answer alone? BaRe-Mem uses verified experience to estimate reliability and guide both decisions.
δ-mem writes earlier context into an 8 × 8 state. δ-Vision reconstructs visual states with lightweight MLPs, reducing computation while retaining every visual token.
A model can resist a harmful question yet lose that safeguard after fine-tuning, or accept work outside its assigned role. CoU, RESTA, WalledEval and OffTopicEval test and address these different failures.
03
Earlier work
Publications
MultimodalityAffective computing2017–2021
Multimodal representation learning
TFN makes interactions between words, facial expressions and voice explicit. MISA separates what these inputs share from what each contributes on its own; Multimodal-InfoMax trains the combined representation to retain information from all three.
MultimodalityAffective computing2017–2019
Emotion in conversations
A sentence's emotion depends on who says it and what came before. MELD supplies labelled conversations with text, speech and video; DialogueRNN tracks each speaker's changing state, while DialogueGCN models the links between their utterances.
MultimodalityAffective computing2016–2019
Sarcasm detection
Our COLING 2016 work used sentiment, emotion and personality cues to recognise sarcasm in tweets. CASCADE added the discussion and the writer's history. MUStARD brought in voice and facial expression, where the joke can be missed by reading the transcript alone.
04
Lab notes
All notesSeptember 2026 · Vernon Toh
ScrambleToolBench: the next step is already in the map
An agent discovers what its tools do. Their names change, and it starts searching again—even when its earlier observations point to the next call.
September 2026 · Vernon Toh
MNIST-PRO: seeing the strokes is not enough
When an agent gets a digit wrong, did it miss the evidence or struggle to put the glimpses together? We use MNIST-PRO to tell these failures apart.
August 2026 · Varun Gumma
IDEAgent: Agentic Quality-Diversity Search
IDEAgent searches for a diverse set of candidate research ideas and records how each candidate changes during refinement.
05
Code, models and data
All releasesconv-emotion
CodeReference implementations of emotion-recognition-in-conversation models, including DialogueRNN.
MELD
DatasetMultimodal multi-party emotion recognition in conversation, used in later audio and multimodal evaluation work.
Tango
ModelAn open latent-diffusion model for text-to-audio, built on an instruction-tuned LLM text encoder rather than a contrastive one.
NORA
ModelA 3B-parameter open generalist vision-language-action model trained on 970k real-world robot demonstrations, built for low inference cost.
δ-mem
ModelOnline associative memory for LLMs — a 0.12% parameter addition that modulates attention through a low-rank interface while the backbone stays frozen.
TangoFlux
ModelText-to-audio generation with flow matching and CRPO preference optimisation. Generates 30 seconds of 44.1 kHz audio in a few seconds on a single A40.
06
People at DeCLaRe
Faculty, research staff and students at NTU Singapore.
07
About DeCLaRe
DeCLaRe, short for Deep Cognition and Language Research, was founded by Soujanya Poria at the Singapore University of Technology and Design in 2019 with Navonil Majumder, Devamanyu Hazarika and Deepanway Ghosal. The lab moved to Nanyang Technological University in 2025.
The robot forms 宣 (xuān), “to declare,” carrying the lab's name directly into its visual identity.










