Hi! 你好We are the DeCLaRe Lab at Nanyang Technological University

Selected research

All research
GE-Act 2.0 project film · 3 min 48 sec · Watch on the original project site ↗

Embodied AIMultimodality

Robot control · 2026

GE-Act 2.0

GE-Act 2.0 learns to predict how a scene will change, then uses that prediction to control a robot. The study tests one checkpoint on 100 real-robot tasks, without task-specific fine-tuning.

An AgiBot–DeCLaRe Lab collaboration, with Renhang Liu as DeCLaRe's lead core contributor and Soujanya Poria contributing as an academic advisor.

Model and experiments

MultimodalityEfficiency

Audio and music

TangoFlux

PromptA basketball bounces rhythmically on a court, shoes squeak against the floor, and a referee’s whistle cuts through the air.

Compare with Tango 2

2016–2021

Earlier work

Publications

Recent highlights

All updates ↗

Lab notes

All notes ↗
A successful MNIST-PRO episode: four glimpses along a handwritten zero, ending with the correct answer at step 9

September 2026 · Vernon Toh

MNIST-PRO: seeing the strokes is not enough

When an agent gets a digit wrong, did it miss the evidence or struggle to put the glimpses together? We use MNIST-PRO to tell these failures apart.

Subscribe by RSS ↗

Code, models and data

More releases ↗

Reference implementations of emotion-recognition-in-conversation models, including DialogueRNN.

MELD

Dataset

Multimodal multi-party emotion recognition in conversation, used in later audio and multimodal evaluation work.

1,081 · ↓ 1,843/moResources ↗

Tango

Model

An open latent-diffusion model for text-to-audio, built on an instruction-tuned LLM text encoder rather than a contrastive one.

1,238 · ↓ 51/moResources ↗

NORA

Model

A 3B-parameter open generalist vision-language-action model trained on 970k real-world robot demonstrations, built for low inference cost.

224 · ↓ 447/moResources ↗

δ-mem

Model

Online associative memory for LLMs — a 0.12% parameter addition that modulates attention through a low-rank interface while the backbone stays frozen.

Text-to-audio generation with flow matching and CRPO preference optimisation. Generates 30 seconds of 44.1 kHz audio in a few seconds on a single A40.

882 · ↓ 616/moResources ↗

People at DeCLaRe

Faculty, research staff and students at NTU Singapore.

  • Soujanya Poria
  • Suorong Yang
  • Rishabh Bhardwaj
  • Navonil Majumder
  • Soumitra Sinhahajari
  • Jingdi Lei

About DeCLaRe

DeCLaRe, short for Deep Cognition and Language Research, was founded by Soujanya Poria at the Singapore University of Technology and Design in 2019 with Navonil Majumder, Devamanyu Hazarika and Deepanway Ghosal. The lab moved to Nanyang Technological University in 2025.

Lab's identity

The robot forms 宣 (xuān), “to declare,” carrying the lab's name directly into its visual identity.

宣 · Xuān The robot is shaped from 宣, meaning “to declare.”