DeCLaRe Lab · NTU Singapore

Research

We study how models understand language, images and sound, learn from experience, and act reliably.

Safety

Trustworthiness

Multimodality

Multimodal representation learning

TFN, MISA and Multimodal-InfoMax study how to combine language, vision and audio: by modelling their interactions, separating shared and modality-specific information, or preserving information during fusion.

MISA detail: a modality-invariant encoder maps language, audio and vision into shared representations

Multimodality2020

MISA

MISA gives language, vision and audio both shared and modality-specific representations, then combines all six vectors to predict sentiment or humour.

MISA · 294 stars
Methods and results

AI for Science

Efficiency

Embodied AI

Code, models and data

  • Audio and music generation · Model · ICLR 2026

    TangoFlux

    Text-to-audio generation with flow matching and CRPO preference optimisation. Generates 30 seconds of 44.1 kHz audio in a few seconds on a single A40.

    882 stars 679/mo

  • Audio and music generation · Model · ACM MM 2023

    Tango

    An open latent-diffusion model for text-to-audio, built on an instruction-tuned LLM text encoder rather than a contrastive one.

    1,238 stars 60/mo

  • Audio and music generation · Model · NAACL 2024

    Mustango

    Text-to-music generation conditioned on explicit musical structure — chords, beats, tempo and key — rather than free-text mood alone.

    395 stars 2,139/mo

  • Audio and music generation · Model · 2025

    JAM

    A compact flow-based song generator with word-level lyric timing control and aesthetic alignment from synthetic preference pairs.

    169 stars 85/mo

  • Audio and music generation · Dataset · 2023

    TangoPromptBank

    The text-audio caption corpus Tango was pretrained on.

    4,497/mo

  • Embodied AI and robotics · Model · 2025

    NORA

    A 3B-parameter open generalist vision-language-action model trained on 970k real-world robot demonstrations, built for low inference cost.

    224 stars 478/mo

  • Embodied AI and robotics · Model · 2025

    NORA-1.5

    NORA with a flow-matching action expert, post-trained against world-model and action-based preference rewards.

    109 stars 21/mo

  • Embodied AI and robotics · Model · 2025

    Emma-X

    An embodied multimodal action model that plans with grounded chain-of-thought and look-ahead spatial reasoning.

    84 stars 52/mo

  • Trustworthiness and safety · Model + data · ICLR 2025

    Trust-Align

    Measures whether RAG answers are supported and fully cited, and trains models to refuse when the retrieved evidence is insufficient.

    76 stars 439/mo

  • Trustworthiness and safety · Code · ACL 2024

    RESTA

    Restoring safety in fine-tuned language models by adding a safety vector back through task arithmetic, without losing downstream capability.

    33 stars

  • Trustworthiness and safety · Code + data · 2023

    Red-Instruct

    Chain-of-Utterances red-teaming, with the RED-EVAL benchmark and the HarmfulQA dataset for probing and repairing harmful behaviour.

    111 stars 1,523/mo

  • Trustworthiness and safety · Benchmark · ICLR 2026

    OffTopicEval

    Measures whether a task-scoped LLM agent accepts what falls inside its remit and refuses what falls outside it.

    11 stars 131/mo

  • Efficiency, memory and agents · Code · ICML 2026

    Data Agent

    Learns which training examples to select as the target model changes, using difficulty and uncertainty as feedback.

    14 stars

  • Efficiency, memory and agents · Model · 2026

    δ-mem

    Online associative memory for LLMs — a 0.12% parameter addition that modulates attention through a low-rank interface while the backbone stays frozen.

    257 stars

  • Efficiency, memory and agents · Code · 2026

    IDEAgent

    Searches for a set of research ideas that meet quality thresholds without repeating one another.

    18 stars

  • Efficiency, memory and agents · Code · 2026

    EFLA

    Error-free linear attention derived as an exact solution from continuous-time dynamics rather than an approximation.

    77 stars

  • Efficiency, memory and agents · Model · 2023

    Flan-Alpaca

    Extending Stanford Alpaca's synthetic instruction tuning to Flan-T5, with the Flan-mini instruction collection released alongside.

    354 stars 1,764/mo

  • Efficiency, memory and agents · Benchmark · 2023

    InstructEval

    Evaluates instruction-tuned models on problem solving, writing and alignment to human values.

    553 stars 55/mo

  • Multimodal understanding · Dataset · ACL 2019

    MELD

    Multimodal multi-party emotion recognition in conversation, used in later audio and multimodal evaluation work.

    1,081 stars 1,830/mo

  • Multimodal understanding · Code · 2019

    conv-emotion

    Reference implementations of emotion-recognition-in-conversation models, including DialogueRNN.

    1,523 stars

  • Multimodal understanding · Code · 2020

    multimodal-deep-learning

    A collection of multimodal representation-learning and fusion models for sentiment and emotion, kept as a single reference implementation set.

    920 stars

  • Multimodal understanding · Code · ACM MM 2020

    MISA

    Decomposes multimodal representations into modality-invariant and modality-specific subspaces before fusion.

    294 stars

  • Multimodal understanding · Dataset · 2021

    RECCON

    Recognising the cause of an emotion in conversation, not only the emotion itself.

    191 stars

  • Multimodal understanding · Dataset · 2022

    CICERO

    Dialogue-level commonsense inference — causes, motivations, reactions and consequences drawn from conversation.

    63 stars 191/mo

  • Multimodal understanding · Benchmark · 2024

    PuzzleVQA

    Abstract visual puzzles and algorithmic reasoning tasks that separate perception from inductive reasoning in multimodal models.

    119 stars 454/mo

  • Reading lists · Reading list · 2018

    Awesome Sentiment Analysis

    537 stars

  • Reading lists · Reading list · 2019

    Awesome ERC

    283 stars

Updated 14 September 2026 · All 80 public GitHub repositories · 51 models and 21 datasets on Hugging Face