DINOv2
Free while signed in. Answers cite the passages they came from.

Meta's self-supervised vision foundation model producing robust features without labels.
Fully self-supervised: Trained purely with SSL on 142M curated images - no labels needed, just clever pretraining objectives.
Universal features: Produces features useful for image classification, instance retrieval, video understanding, depth estimation, and pixel-level tasks.
Frozen-backbone usage: Features work well with simple linear probes, no fine-tuning - making DINOv2 a drop-in visual backbone.
Vision foundation standard: Became the default vision backbone for open-source VLMs (LLaVA, InternVL) and vision research through 2024.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack