RT-2
Free while signed in. Answers cite the passages they came from.

Google DeepMind's end-to-end vision-language-action model that learns from both web and robotics data to control robots.
VLA architecture: Treats robot actions as another language the model generates - actions are tokenized and output in the same stream as text tokens.
Web-scale knowledge transfer: Leverages internet-scale VLM pretraining so the robot can reason about novel objects and symbols it never saw in robotics data (e.g., "pick up the extinct animal").
Emergent semantic reasoning: Shows emergent capabilities like chain-of-thought robotic reasoning and multi-stage task planning absent in prior RT-1.
Robot foundation models: Established the VLA paradigm that dominated 2024 robotics research (OpenVLA, RT-X, π0) and moved robotics firmly into the foundation-model era.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack