🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
Training

V-JEPA

Free while signed in. Answers cite the passages they came from.

Paper preview
V-JEPA
The curator’s take

Meta's V-JEPA learns visual representations by predicting features in masked video regions, without pretrained image encoders, text, negatives, or reconstruction.

Key points
01

Feature-prediction objective: A student encoder predicts target-encoder features for masked video patches - a pure SSL recipe that bypasses pixel reconstruction and contrastive negatives entirely.

02

2M-video training set: Trained on 2M publicly available videos, big enough to learn rich spatiotemporal features without needing labels or captions.

03

Frozen-evaluation results: A ViT-H/16 V-JEPA hits 81.9% on Kinetics-400, 72.2% on Something-Something-v2, and 77.9% on ImageNet1K in frozen-feature evaluations.

04

Versatile features: The same representation works for motion-heavy (Something-Something) and appearance-heavy (ImageNet) tasks, supporting the JEPA claim that feature prediction produces more general visual features than reconstruction.

Every Monday
Get next week’s papers.

The same picks and the same summaries, in your inbox. Free, and 176 issues deep.

Subscribe on Substack