🚀NEW LABGetting Started with Claude AgentsStart lab
Data

Whole-Body Conditioned Egocentric Video Prediction

First page
Whole-Body Conditioned Egocentric Video Prediction
Paper summary

This paper introduces PEVA, a conditional diffusion transformer that predicts egocentric video conditioned on 3D human body motion. Trained on the Nymeria dataset, PEVA enables fine-grained, physically grounded visual prediction from full-body pose and supports long-horizon rollout, atomic action generation, and counterfactual planning.

Ask this paper

Every Monday
Get next week’s papers.
Subscribe on Substack