OpenEQA
Free while signed in. Answers cite the passages they came from.

Meta's OpenEQA is an open-vocabulary benchmark for embodied question answering: 1,600+ human-written questions across 180+ real-world environments, with a calibrated LLM-as-judge metric that tracks human agreement closely.
Benchmark setup: Questions cover episodic-memory use cases (smart glasses) and active-exploration use cases (mobile robots), demanding that agents reason about the environment they occupy rather than just a single image.
LLM-as-judge scoring: The paper introduces an automatic LLM-powered evaluation protocol with strong correlation to human judgment, solving the open-vocabulary scoring problem that blocks most EQA benchmarks from scaling.
Frontier-model performance: GPT-4V, Claude 3, and Gemini Pro significantly outperform text-only baselines, but their gains come mostly from object recognition - for several question categories, they barely beat a blind-LLM baseline.
Gap to humans: Across the board, state-of-the-art multimodal LLMs lag well behind human performance on OpenEQA, positioning the benchmark as a concrete target for embodied-agent research.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack