V-JEPA

Meta's V-JEPA learns visual representations by predicting features in masked video regions, without pretrained image encoders, text, negatives, or reconstruction.
Ask this paper
Feature-prediction objective: A student encoder predicts target-encoder features for masked video patches - a pure SSL recipe that bypasses pixel reconstruction and contrastive negatives entirely.
2M-video training set: Trained on 2M publicly available videos, big enough to learn rich spatiotemporal features without needing labels or captions.
Frozen-evaluation results: A ViT-H/16 V-JEPA hits 81.9% on Kinetics-400, 72.2% on Something-Something-v2, and 77.9% on ImageNet1K in frozen-feature evaluations.
Versatile features: The same representation works for motion-heavy (Something-Something) and appearance-heavy (ImageNet) tasks, supporting the JEPA claim that feature prediction produces more general visual features than reconstruction.