🚀NEW LABGetting Started with Claude AgentsStart lab
Multimodal

Hallucination in LVLMs

First page
Hallucination in LVLMs
Paper summary

A survey specifically scoped to hallucination in Large Vision-Language Models, a phenomenon that differs substantially from text-only LLM hallucination.

Ask this paper

Key points
01

Hallucination taxonomy: Distinguishes object, attribute, and relation hallucinations in LVLMs and discusses how each arises from different architectural and data pressures.

02

Causes: Traces hallucinations to biased training data, weak visual grounding, language prior dominance, and weaknesses in the vision encoder.

03

Evaluation: Reviews benchmarks (POPE, MMHal-Bench, HallusionBench) and automatic evaluation metrics specific to visual hallucination.

04

Mitigation: Catalogs mitigation strategies including improved data curation, better visual alignment, decoding-time interventions, and RLHF-style preference optimization.

Every Monday
Get next week’s papers.
Subscribe on Substack