- Safety Will task-specific agents stay within scope?OffTopicEval Does safety survive fine-tuning?RESTA
- Trustworthiness When should agents trust one another?Epistemic Context Learning Do RAG citations support each claim?Trust-Align
- Multimodality Can models infer rules from visual puzzles?Puzzle Prodigies Can models reason directly from pixels?MIRAS
- AI for Science Can agents find diverse research ideas?IDEAgent Do LLMs judge scientific novelty reliably?LLM-as-Judge for Novelty
- Efficiency Can online memory remain compact?delta-mem Can agents learn continually from memory?Sigma-Mem
- Embodied AI Can preference training improve VLA policies?NORA 1.5 Which data bottlenecks limit VLAs?10 Open Challenges for VLAs
Research
Selected Projects
2017–2023
Landmark work
Multimodality · 2017
Tensor Fusion Network (TFN)
TFN represents every unimodal, bimodal and trimodal interaction among language, facial behaviour and voice with one outer product. The paper tests whether that explicit structure improves sentiment prediction over concatenation.
Multimodality · 2018–2019
MELD + DialogueRNN
MELD pairs 13,708 conversational utterances with text, audio, video, emotion and sentiment labels. DialogueRNN tests whether tracking each speaker separately improves emotion classification.
Multimodality · Efficiency · 2023–2026
TANGO
Across five systems, this work moves from conditioning short sound generation on language, to learning event order from preference pairs, exposing musical controls, reducing sampling time and placing individual words inside full songs.
2025–2026
Current directions
Multimodality · Efficiency · 2026
ScrambleToolBench + MNIST-PRO
ScrambleToolBench measures whether agents can identify and reuse obfuscated tools as their mappings change. MNIST-PRO measures whether vision-language models can assemble movable image glimpses into a correct spatial state.
Efficiency · Trustworthiness · 2026
δ-mem + Σ-Mem
δ-mem writes prior context into a fixed 8 × 8 state that steers a frozen language model. Σ-Mem records task-conditioned peer correctness and co-error patterns for multi-agent selection, routing and voting.
Trustworthiness · Safety · 2025–2026
KAIROS + ECL
KAIROS varies prior rapport, current peer behaviour and model confidence across 3,000 questions. Epistemic Context Learning makes an agent infer who has been reliable before it sees their advice on the next question.
Lab Notes
IDEAgent: Agentic Quality-Diversity Search
IDEAgent searches for a diverse set of candidate research ideas and records how each candidate changes during refinement.
Toward Efficient Data-Centric Training, Part II
PODS changes how much data is used at each stage of training instead of fixing one selection ratio.
Toward Efficient Data-Centric Training, Part I
Data Agent updates its data-selection policy as the target model trains.
Code, Models and Data
Reference implementations of emotion-recognition-in-conversation models, including DialogueRNN.
1,523 MELD DatasetMultimodal multi-party emotion recognition in conversation, used in later audio and multimodal evaluation work.
1,0791,944/mo Tango ModelAn open latent-diffusion model for text-to-audio, built on an instruction-tuned LLM text encoder rather than a contrastive one.
1,23850/mo NORA ModelA 3B-parameter open generalist vision-language-action model trained on 970k real-world robot demonstrations, built for low inference cost.
224456/mo δ-mem ModelOnline associative memory for LLMs — a 0.12% parameter addition that modulates attention through a low-rank interface while the backbone stays frozen.
255 TangoFlux ModelText-to-audio generation with flow matching and CRPO preference optimisation. Generates 30 seconds of 44.1 kHz audio in a few seconds on a single A40.
882640/moRecent highlights
ScrambleToolBench: adapting to hidden tool changes
A terminal benchmark for agents that must infer undocumented tool behaviour and recover when the tool-to-function mapping changes.
Research internships for DeCLaRe students
Four students started internships at Tencent, MiniMax and Meta.
δ-mem featured in VentureBeat
The article covers δ-mem's compact online state and its low-rank interface to a frozen language model.
Embodied Foundational Models funded by CNRS@CREATE and NRF
A 2026–2029 project on generalist vision-language-action models for embodied AI.
- Soujanya Poria
- Suorong Yang
- Rishabh Bhardwaj
- Navonil Majumder
- Soumitra Sinhahajari
- Jingdi Lei
About DeCLaRe
DeCLaRe, short for Deep Cognition and Language Research, was founded by Soujanya Poria at the Singapore University of Technology and Design in 2019 with Navonil Majumder, Devamanyu Hazarika and Deepanway Ghosal. The lab moved to Nanyang Technological University in 2025.
The robot forms 宣 (xuān), “to declare,” carrying the lab's name directly into its visual identity.