Hallucination in LVLMs
Free while signed in. Answers cite the passages they came from.

A survey specifically scoped to hallucination in Large Vision-Language Models, a phenomenon that differs substantially from text-only LLM hallucination.
Hallucination taxonomy: Distinguishes object, attribute, and relation hallucinations in LVLMs and discusses how each arises from different architectural and data pressures.
Causes: Traces hallucinations to biased training data, weak visual grounding, language prior dominance, and weaknesses in the vision encoder.
Evaluation: Reviews benchmarks (POPE, MMHal-Bench, HallusionBench) and automatic evaluation metrics specific to visual hallucination.
Mitigation: Catalogs mitigation strategies including improved data curation, better visual alignment, decoding-time interventions, and RLHF-style preference optimization.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack