Dynalang (Agents Model the World with Language)
Free while signed in. Answers cite the passages they came from.

UC Berkeley's Dynalang agent learns a multimodal world model predicting future text, video, and rewards.
Multimodal world model: Jointly predicts future language, video, and rewards, treating language as another stream of observation/prediction rather than just policy input.
Instruction-following: Learns to follow instructions in visually and linguistically complex domains, grounded in the world model's predictions.
Cross-domain applicability: Applied to multiple embodied environments, showing the language-inclusive world-model approach is general.
Research direction: Foreshadows the "video-plus-language world model" direction that would grow prominent in 2024 (e.g., Sora's world simulator framing).
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack