DeCLaRe Lab · NTU Singapore

Updates

In the news

2026

Paper

ScrambleToolBench: adapting to hidden tool changes

A terminal benchmark for agents that must infer undocumented tool behaviour and recover when the tool-to-function mapping changes.

Paper

Σ-Mem: online reliability memory for multi-agent systems

Σ-Mem updates a symmetric record of pairwise reliability from post-decision feedback without retraining the agents.

Paper

IDEAgent: quality-diversity search for research ideas

IDEAgent searches for a set of ideas that meet quality thresholds without repeating one another, and tracks each candidate through repair and refinement.

Paper

RQ-Bench: testing LLM judgements of scientific novelty

RQ-Bench derives research questions from published papers and authors' accounts, then compares LLM novelty scores with expert judgements.

Paper

GRAIL: token-level advantages for RL reasoning

GRAIL uses gradient-activation saliency to assign token-level advantage weights in reinforcement learning with verifiable rewards.

Paper

Two papers accepted to ICML 2026

The papers study adaptive data selection during training and search-guided reasoning over video.

  • Data Agent: end-to-end dynamic data selection for training-aware efficiency
  • Chain-of-Glimpse: search-guided, object-grounded progressive reasoning over video
Paper

PODS: scheduling data volume during training

PODS alternates low- and high-data phases under a fixed cumulative budget and can be used with existing data-selection methods.

Paper

Stacked from One: longer context through self-injection

A multi-scale self-injection method for extending usable context without retraining the base model from scratch.

Paper

From Perception to Action: interactive visual reasoning

A benchmark for spatial and visual reasoning during interaction with simulated physical environments.

Paper

Four papers accepted to ICLR 2026

The papers cover text-to-audio generation, task boundaries, multi-agent social influence and retrieval evaluation.

  • TangoFlux: fast, faithful text-to-audio generation with flow matching
  • OffTopicEval: whether task-specific agents accept in-domain queries while refusing the rest
  • LLMs Can't Handle Peer Pressure: susceptibility to multi-agent social influence
  • Demystifying Deep Search: a hint-free multi-hop evaluation of search agents
Paper

Epistemic Context Learning: evidence and peer reliability

Studies how LLM agents update beliefs about one another from evidence gathered during multi-agent interaction.

Grant

Embodied Foundational Models funded by CNRS@CREATE and NRF

A 2026–2029 project on generalist vision-language-action models for embodied AI.

Grant

Toward Generalist Vision Language Action Models funded by KLASS

A 2026–2028 project on action grounding and evaluation across robot platforms.

Award

Highly Cited Researcher recognition

Soujanya Poria was recognized by Web of Science as a Highly Cited Researcher.

Grant

Google DeepMind GCP grant

S$100K in cloud compute for language, multimodal and agent research.

2025

Paper

Two papers accepted to AAAI 2026

A position paper on future directions in vision-language-action research and a study of multimodal reasoning on visual puzzles.

  • 10 Open Challenges Steering the Future of Vision-Language-Action Models
  • Tracking the Evolution of Multimodal Reasoning on Visual Puzzles
Paper

NORA-1.5: preference post-training for vision-language-action models

NORA-1.5 adds a flow-matching action expert and preference pairs scored by a world model and trajectory deviation.

Paper

10 Open Challenges Steering the Future of Vision-Language-Action Models

A position paper on spatial reasoning, world dynamics, post-training and cross-embodiment action generalisation for VLA models.

Paper

Demystifying Deep Search: hint-free multi-hop evaluation

WebDetective tests whether web agents can find multi-hop reasoning chains without hints embedded in the question.

Paper

OffTopicEval: out-of-domain refusal in LLM agents

Tests whether task-specific agents accept requests within scope and refuse requests outside it across 20 open-weight models.

Paper

Training Vision-Language PRMs for Test-Time Scaling in Multimodal Reasoning

Studies dataset synthesis, perception-level supervision and test-time scaling for vision-language process reward models.

Paper

LLMs Can't Handle Peer Pressure

KAIROS tests how rapport, peer actions and model confidence affect consensus in multi-agent LLM systems.

Lab

DeCLaRe Lab moves to Nanyang Technological University

DeCLaRe moved from SUTD to NTU Singapore in August 2025.

Grant

Meta Audiobox Research Grant

Research funding for audio generation and multimodal generative modelling.

Release

NORA and NORA-1.5 released

The lab released code and model weights for both vision-language-action models.

Paper

Trust-Score/Trust-Align and MOOSE-Chem accepted at ICLR

Papers studying trustworthy retrieval-augmented generation and chemistry hypothesis discovery.