Foundation Models in Vision
Free while signed in. Answers cite the passages they came from.

A comprehensive survey on foundational models for computer vision and their open research directions.
Landscape mapping: Reviews textually prompted (CLIP, ALIGN), visually prompted (SAM), and generative (DALL-E, Imagen) vision foundation models in one unified taxonomy.
Challenges enumerated: Identifies open problems in evaluation, grounding, hallucination, compositionality, and domain-specific adaptation for CV.
Cross-modal trends: Analyzes how vision foundation models increasingly borrow from LLM training recipes (instruction tuning, RLHF).
Reference for researchers: Became a go-to survey for new researchers entering vision foundation-model research in late 2023.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack