Dynalang (Agents Model the World with Language)
First page

Paper summary
UC Berkeley's Dynalang agent learns a multimodal world model predicting future text, video, and rewards.
Ask this paper
01
Multimodal world model: Jointly predicts future language, video, and rewards, treating language as another stream of observation/prediction rather than just policy input.
02
Instruction-following: Learns to follow instructions in visually and linguistically complex domains, grounded in the world model's predictions.
03
Cross-domain applicability: Applied to multiple embodied environments, showing the language-inclusive world-model approach is general.
04
Research direction: Foreshadows the "video-plus-language world model" direction that would grow prominent in 2024 (e.g., Sora's world simulator framing).