Deep Cognition and Language ResearchNTU Singapore

We study how models connect language with speech, sound, images, video and robot action.

01

Research themes

  1. 01

    Safety

    Safety research

  2. 02

    Trustworthiness

    Trustworthiness research

  3. 03

    Multimodality

    Multimodality research

  4. 04

    AI for Science

    AI for Science research

  5. 05

    Efficiency

    Efficiency research

  6. 06

    Embodied AI

    Embodied AI research

02

Selected recent research

All research
Narrated overview · 3 min 12 sec · Transcript

Embodied AIEfficiency2026

MemBodied: memory for what a robot can no longer see

A robot's next action can depend on what is no longer visible. MemBodied remembers actions and their observed outcomes in a fixed-size recurrent state.

MemBodied reads from associative memory before acting, then writes the action together with its observed outcome. A separate anchor preserves the initial scene.

Across five RMBench tasks it reaches 50.0% mean success, against 6.4% for the stateless baseline.

MultimodalityEfficiency2023–2026

TangoFlux

PromptA basketball bounces rhythmically on a court, shoes squeak against the floor, and a referee’s whistle cuts through the air.

Compare with Tango 2

Embodied AIMultimodality2025

NORA and NORA-1.5

NORA learns to turn images and instructions into robot actions from 970,000 demonstrations. NORA-1.5 adds preference training: a world model predicts the outcomes of candidate actions, which are then ranked alongside their distance from demonstrated actions.

Models and robot trials

03

Earlier work

Publications

04

Lab notes

All notes

September 2026 · Vernon Toh

MNIST-PRO: seeing the strokes is not enough

When an agent gets a digit wrong, did it miss the evidence or struggle to put the glimpses together? We use MNIST-PRO to tell these failures apart.

05

Code, models and data

All releases

conv-emotion

Code

Reference implementations of emotion-recognition-in-conversation models, including DialogueRNN.

1,527

MELD

Dataset

Multimodal multi-party emotion recognition in conversation, used in later audio and multimodal evaluation work.

1,088↓ 2,082/mo

Tango

Model

An open latent-diffusion model for text-to-audio, built on an instruction-tuned LLM text encoder rather than a contrastive one.

1,240↓ 67/mo

NORA

Model

A 3B-parameter open generalist vision-language-action model trained on 970k real-world robot demonstrations, built for low inference cost.

224↓ 611/mo

δ-mem

Model

Online associative memory for LLMs — a 0.12% parameter addition that modulates attention through a low-rank interface while the backbone stays frozen.

261

TangoFlux

Model

Text-to-audio generation with flow matching and CRPO preference optimisation. Generates 30 seconds of 44.1 kHz audio in a few seconds on a single A40.

886↓ 877/mo

06

People at DeCLaRe

Faculty, research staff and students at NTU Singapore.

  • Soujanya Poria
  • Suorong Yang
  • Rishabh Bhardwaj
  • Navonil Majumder
  • Jingdi Lei
  • Keith Kwok

07

About DeCLaRe

DeCLaRe, short for Deep Cognition and Language Research, was founded by Soujanya Poria at the Singapore University of Technology and Design in 2019 with Navonil Majumder, Devamanyu Hazarika and Deepanway Ghosal. The lab moved to Nanyang Technological University in 2025.

The robot forms 宣 (xuān), “to declare,” carrying the lab's name directly into its visual identity.