Research

Safety

Trustworthiness

Multimodality

Multimodal representation learning

TFN, MISA and Multimodal-InfoMax study how to combine language, vision and audio: by modelling their interactions, separating shared and modality-specific information, or preserving information during fusion.

AI for Science

Efficiency

Embodied AI

Code, models and data

  • Audio and music generation · Model · ICLR 2026

    TangoFlux

    Text-to-audio generation with flow matching and CRPO preference optimisation. Generates 30 seconds of 44.1 kHz audio in a few seconds on a single A40.

    886 stars 877/mo

  • Audio and music generation · Model · ACM MM 2023

    Tango

    An open latent-diffusion model for text-to-audio, built on an instruction-tuned LLM text encoder rather than a contrastive one.

    1,240 stars 67/mo

  • Audio and music generation · Model · NAACL 2024

    Mustango

    Text-to-music generation conditioned on explicit musical structure — chords, beats, tempo and key — rather than free-text mood alone.

    396 stars 1,778/mo

  • Audio and music generation · Model · 2025

    JAM

    A compact flow-based song generator with word-level lyric timing control and aesthetic alignment from synthetic preference pairs.

    170 stars 100/mo

  • Audio and music generation · Dataset · 2023

    TangoPromptBank

    The text-audio caption corpus Tango was pretrained on.

    320/mo

  • Embodied AI and robotics · Model · 2025

    NORA

    A 3B-parameter open generalist vision-language-action model trained on 970k real-world robot demonstrations, built for low inference cost.

    224 stars 611/mo

  • Embodied AI and robotics · Model · 2025

    NORA-1.5

    NORA with a flow-matching action expert, post-trained against world-model and action-based preference rewards.

    109 stars 38/mo

  • Embodied AI and robotics · Model · 2025

    Emma-X

    An embodied multimodal action model that plans with grounded chain-of-thought and look-ahead spatial reasoning.

    84 stars 80/mo

  • Trustworthiness and safety · Model + data · ICLR 2025

    Trust-Align

    Measures whether RAG answers are supported and fully cited, and trains models to refuse when the retrieved evidence is insufficient.

    76 stars 196/mo

  • Trustworthiness and safety · Code · ACL 2024

    RESTA

    Restoring safety in fine-tuned language models by adding a safety vector back through task arithmetic, without losing downstream capability.

    33 stars

  • Trustworthiness and safety · Code + data · 2023

    Red-Instruct

    Chain-of-Utterances red-teaming, with the RED-EVAL benchmark and the HarmfulQA dataset for probing and repairing harmful behaviour.

    111 stars 1,943/mo

  • Trustworthiness and safety · Benchmark · ICLR 2026

    OffTopicEval

    Measures whether a task-scoped LLM agent accepts what falls inside its remit and refuses what falls outside it.

    12 stars 130/mo

  • Efficiency, memory and agents · Code · ICML 2026

    Data Agent

    Learns which training examples to select as the target model changes, using difficulty and uncertainty as feedback.

    14 stars

  • Efficiency, memory and agents · Model · 2026

    δ-mem

    Online associative memory for LLMs — a 0.12% parameter addition that modulates attention through a low-rank interface while the backbone stays frozen.

    261 stars

  • Efficiency, memory and agents · Code · 2026

    IDEAgent

    Searches for a set of research ideas that meet quality thresholds without repeating one another.

    19 stars

  • Efficiency, memory and agents · Code · 2026

    EFLA

    Error-free linear attention derived as an exact solution from continuous-time dynamics rather than an approximation.

    77 stars

  • Efficiency, memory and agents · Model · 2023

    Flan-Alpaca

    Extending Stanford Alpaca's synthetic instruction tuning to Flan-T5, with the Flan-mini instruction collection released alongside.

    354 stars 2,065/mo

  • Efficiency, memory and agents · Benchmark · 2023

    InstructEval

    Evaluates instruction-tuned models on problem solving, writing and alignment to human values.

    553 stars 48/mo

  • Multimodal understanding · Dataset · ACL 2019

    MELD

    Multimodal multi-party emotion recognition in conversation, used in later audio and multimodal evaluation work.

    1,088 stars 2,082/mo

  • Multimodal understanding · Code · 2019

    conv-emotion

    Reference implementations of emotion-recognition-in-conversation models, including DialogueRNN.

    1,527 stars

  • Multimodal understanding · Code · 2020

    multimodal-deep-learning

    A collection of multimodal representation-learning and fusion models for sentiment and emotion, kept as a single reference implementation set.

    919 stars

  • Multimodal understanding · Code · ACM MM 2020

    MISA

    Decomposes multimodal representations into modality-invariant and modality-specific subspaces before fusion.

    296 stars

  • Multimodal understanding · Dataset · 2021

    RECCON

    Recognising the cause of an emotion in conversation, not only the emotion itself.

    191 stars

  • Multimodal understanding · Dataset · 2022

    CICERO

    Dialogue-level commonsense inference — causes, motivations, reactions and consequences drawn from conversation.

    63 stars 175/mo

  • Multimodal understanding · Benchmark · 2024

    PuzzleVQA

    Abstract visual puzzles and algorithmic reasoning tasks that separate perception from inductive reasoning in multimodal models.

    119 stars 427/mo

  • Reading lists · Reading list · 2018

    Awesome Sentiment Analysis

    538 stars

  • Reading lists · Reading list · 2019

    Awesome ERC

    284 stars

Updated 28 September 2026 · All 81 public GitHub repositories · 51 models and 21 datasets on Hugging Face