Evaluating Verifiability in Generative Search Engines
First page

Paper summary
Audits popular generative search engines for citation accuracy.
Ask this paper
01
Human evaluation: Performs rigorous human evaluation of Bing Chat, Perplexity AI, and NeevaAI responses.
02
Citation failure rate: Finds only 52% of generated sentences are supported by citations and only 75% of citations actually support the claim.
03
Verifiability gap: Reveals a significant gap between generative search engines' citation promises and their actual reliability.
04
Trust-in-AI research: Important empirical foundation for subsequent research on grounded generation and RAG accuracy.