Depth Anything

A robust monocular depth estimator designed to handle "any image under any circumstance" by scaling self-training on unlabeled data rather than hunting for bigger labeled sets.
Ask this paper
62M unlabeled images: The data engine automatically annotates ~62M unlabeled images using a teacher model, then uses these pseudo-labels for student training - a classic but here-industrialized recipe.
Stronger supervision signals: Introduces auxiliary supervision that forces the student to inherit semantic priors from a pretrained encoder, preventing the usual failure modes of naive self-training at this scale.
SoTA with fine-tuning: Beyond strong zero-shot generalization, fine-tuning on downstream depth datasets sets new state-of-the-art on standard benchmarks.
Enhanced ControlNet: Depth-conditioned ControlNet built on Depth Anything produces noticeably cleaner depth-guided image generation, highlighting downstream impact beyond perception tasks.