DINOv2
First page

Paper summary
Meta's self-supervised vision foundation model producing robust features without labels.
Ask this paper
01
Fully self-supervised: Trained purely with SSL on 142M curated images - no labels needed, just clever pretraining objectives.
02
Universal features: Produces features useful for image classification, instance retrieval, video understanding, depth estimation, and pixel-level tasks.
03
Frozen-backbone usage: Features work well with simple linear probes, no fine-tuning - making DINOv2 a drop-in visual backbone.
04
Vision foundation standard: Became the default vision backbone for open-source VLMs (LLaVA, InternVL) and vision research through 2024.