🚀NEW LABGetting Started with Claude AgentsStart lab
Training

V-JEPA

Paper preview
V-JEPA
Paper summary

Meta's V-JEPA learns visual representations by predicting features in masked video regions, without pretrained image encoders, text, negatives, or reconstruction.

Ask this paper

Key points
01

Feature-prediction objective: A student encoder predicts target-encoder features for masked video patches - a pure SSL recipe that bypasses pixel reconstruction and contrastive negatives entirely.

02

2M-video training set: Trained on 2M publicly available videos, big enough to learn rich spatiotemporal features without needing labels or captions.

03

Frozen-evaluation results: A ViT-H/16 V-JEPA hits 81.9% on Kinetics-400, 72.2% on Something-Something-v2, and 77.9% on ImageNet1K in frozen-feature evaluations.

04

Versatile features: The same representation works for motion-heavy (Something-Something) and appearance-heavy (ImageNet) tasks, supporting the JEPA claim that feature prediction produces more general visual features than reconstruction.

Every Monday
Get next week’s papers.
Subscribe on Substack