🚀NEW LABGetting Started with Claude AgentsStart lab
Training

Depth Anything

First page
Depth Anything
Paper summary

A robust monocular depth estimator designed to handle "any image under any circumstance" by scaling self-training on unlabeled data rather than hunting for bigger labeled sets.

Ask this paper

Key points
01

62M unlabeled images: The data engine automatically annotates ~62M unlabeled images using a teacher model, then uses these pseudo-labels for student training - a classic but here-industrialized recipe.

02

Stronger supervision signals: Introduces auxiliary supervision that forces the student to inherit semantic priors from a pretrained encoder, preventing the usual failure modes of naive self-training at this scale.

03

SoTA with fine-tuning: Beyond strong zero-shot generalization, fine-tuning on downstream depth datasets sets new state-of-the-art on standard benchmarks.

04

Enhanced ControlNet: Depth-conditioned ControlNet built on Depth Anything produces noticeably cleaner depth-guided image generation, highlighting downstream impact beyond perception tasks.

Every Monday
Get next week’s papers.
Subscribe on Substack