Foundation Models in Vision
First page

Paper summary
A comprehensive survey on foundational models for computer vision and their open research directions.
Ask this paper
01
Landscape mapping: Reviews textually prompted (CLIP, ALIGN), visually prompted (SAM), and generative (DALL-E, Imagen) vision foundation models in one unified taxonomy.
02
Challenges enumerated: Identifies open problems in evaluation, grounding, hallucination, compositionality, and domain-specific adaptation for CV.
03
Cross-modal trends: Analyzes how vision foundation models increasingly borrow from LLM training recipes (instruction tuning, RLHF).
04
Reference for researchers: Became a go-to survey for new researchers entering vision foundation-model research in late 2023.