🚀NEW LABGetting Started with Claude AgentsStart lab
Robotics

RT-2

First page
RT-2
Paper summary

Google DeepMind's end-to-end vision-language-action model that learns from both web and robotics data to control robots.

Ask this paper

Key points
01

VLA architecture: Treats robot actions as another language the model generates - actions are tokenized and output in the same stream as text tokens.

02

Web-scale knowledge transfer: Leverages internet-scale VLM pretraining so the robot can reason about novel objects and symbols it never saw in robotics data (e.g., "pick up the extinct animal").

03

Emergent semantic reasoning: Shows emergent capabilities like chain-of-thought robotic reasoning and multi-stage task planning absent in prior RT-1.

04

Robot foundation models: Established the VLA paradigm that dominated 2024 robotics research (OpenVLA, RT-X, π0) and moved robotics firmly into the foundation-model era.

Every Monday
Get next week’s papers.
Subscribe on Substack