NORA

An open 3B vision-language-action model for robot manipulation

Chia-Yu Hung, Qi Sun, Pengfei Hong, Amir Zadeh, Chuan Li, U-Xuan Tan, Navonil Majumder, Soujanya Poria
DeCLaRe Lab SUTD Lambda Labs

Abstract

NORA is a 3B-parameter vision-language-action model based on Qwen2.5-VL-3B. It is trained on 970,000 real-world robot demonstrations and uses the FAST+ tokenizer to represent action sequences. The smaller backbone is intended to reduce the cost of fine-tuning and inference.

The paper evaluates NORA on robot manipulation and compares its task performance and computational cost with larger VLA models.

WidowX results

NORA evaluation summary across robot tasks

Model

NORA vision-language-action model architecture

Tasks with distractors

With object distraction: Put the pink toy in pot
With human distraction: Put the carrot in pot
With human + object distraction: Put the pink toy in pot

More robot demonstrations

Put the blue cube on the plate
Put the corn and carrot in pan
Put banana and carrot in pot
Move the banana close to the pan
Put carrot in the pot
Put the pink toy at the right corner
Put the carrot and hotdog in pot
Put the blue cube on the plate
Put banana in pot
Put the red bottle and the hamburger in the pan