V-JEPA
Free while signed in. Answers cite the passages they came from.

Meta's V-JEPA learns visual representations by predicting features in masked video regions, without pretrained image encoders, text, negatives, or reconstruction.
Feature-prediction objective: A student encoder predicts target-encoder features for masked video patches - a pure SSL recipe that bypasses pixel reconstruction and contrastive negatives entirely.
2M-video training set: Trained on 2M publicly available videos, big enough to learn rich spatiotemporal features without needing labels or captions.
Frozen-evaluation results: A ViT-H/16 V-JEPA hits 81.9% on Kinetics-400, 72.2% on Something-Something-v2, and 77.9% on ImageNet1K in frozen-feature evaluations.
Versatile features: The same representation works for motion-heavy (Something-Something) and appearance-heavy (ImageNet) tasks, supporting the JEPA claim that feature prediction produces more general visual features than reconstruction.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack