Deep Cognition and Language ResearchNTU Singapore

We study how models connect language with speech, sound, images, video and robot action.

Research themes

  1. Safety

    Safety research

  2. Trustworthiness

    Trustworthiness research

  3. Multimodality

    Multimodality research

  4. AI for Science

    AI for Science research

  5. Efficiency

    Efficiency research

  6. Embodied AI

    Embodied AI research

Selected recent research

All research

EfficiencyNeurIPS 2026 · Spotlight

EFLA

Exact Flow Linear Attention replaces the approximate delta-rule update with its exact continuous-time solution, preserving linear-time attention and adding no parameters.

Ten task demonstrations · Silent film · 3 min 41 sec · Task goals and subtasks

Embodied AIOctober 2026

RoboQuest: agents that search, inspect and test

Introducing RoboQuest, a benchmark of ten mobile manipulation tasks where a robot must act to discover information hidden at the start.

Find targets, turn objects to inspect hidden marks, and test unfamiliar mechanisms. The best of five frontier multimodal agents succeeds in 23% of episodes.

Narrated overview · 3 min 12 sec · Real-robot trials

Embodied AIEfficiency2026

MemBodied: memory for what a robot can no longer see

A robot's next action can depend on what is no longer visible. MemBodied remembers actions and their observed outcomes in a fixed-size recurrent state.

MemBodied reads from associative memory before acting, then writes the action together with its observed outcome. A separate anchor preserves the initial scene.

Across five RMBench tasks it reaches 50.0% mean success, against 6.4% for the stateless baseline.

MultimodalityEfficiency2023–2026

TangoFlux

PromptA basketball bounces rhythmically on a court, shoes squeak against the floor, and a referee’s whistle cuts through the air.

Compare with Tango 2

Embodied AIMultimodality2025

NORA and NORA-1.5

NORA learns to turn images and instructions into robot actions from 970,000 demonstrations. NORA-1.5 adds preference training: a world model predicts the outcomes of candidate actions, which are then ranked alongside their distance from demonstrated actions.

Models and robot trials

Earlier work

Publications

Lab notes

All notes

September 2026 · Vernon Toh

MNIST-PRO: seeing the strokes is not enough

When an agent gets a digit wrong, did it miss the evidence or struggle to put the glimpses together? We use MNIST-PRO to tell these failures apart.

Code, models and data

All releases

conv-emotion

Code

Reference implementations of emotion-recognition-in-conversation models, including DialogueRNN.

1,528

MELD

Dataset

Multimodal multi-party emotion recognition in conversation, used in later audio and multimodal evaluation work.

1,090↓ 2,303/mo

Tango

Model

An open latent-diffusion model for text-to-audio, built on an instruction-tuned LLM text encoder rather than a contrastive one.

1,241↓ 57/mo

NORA

Model

A 3B-parameter open generalist vision-language-action model trained on 970k real-world robot demonstrations, built for low inference cost.

224↓ 569/mo

δ-mem

Model

Online associative memory for LLMs — a 0.12% parameter addition that modulates attention through a low-rank interface while the backbone stays frozen.

263

TangoFlux

Model

Text-to-audio generation with flow matching and CRPO preference optimisation. Generates 30 seconds of 44.1 kHz audio in a few seconds on a single A40.

886↓ 1,060/mo

People at DeCLaRe

  • Soujanya Poria
  • Suorong Yang
  • Rishabh Bhardwaj
  • Navonil Majumder
  • Jingdi Lei
  • Keith Kwok

About DeCLaRe

DeCLaRe, short for Deep Cognition and Language Research, was founded by Soujanya Poria at the Singapore University of Technology and Design in 2019 with Navonil Majumder, Devamanyu Hazarika and Deepanway Ghosal. The lab moved to Nanyang Technological University in 2025.

The robot forms 宣 (xuān), “to declare”.