🚀NEW LABGetting Started with Claude AgentsStart lab
Agents · Multimodal

Dynalang (Agents Model the World with Language)

First page
Dynalang (Agents Model the World with Language)
Paper summary

UC Berkeley's Dynalang agent learns a multimodal world model predicting future text, video, and rewards.

Ask this paper

Key points
01

Multimodal world model: Jointly predicts future language, video, and rewards, treating language as another stream of observation/prediction rather than just policy input.

02

Instruction-following: Learns to follow instructions in visually and linguistically complex domains, grounded in the world model's predictions.

03

Cross-domain applicability: Applied to multiple embodied environments, showing the language-inclusive world-model approach is general.

04

Research direction: Foreshadows the "video-plus-language world model" direction that would grow prominent in 2024 (e.g., Sora's world simulator framing).

Every Monday
Get next week’s papers.
Subscribe on Substack