🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
Evaluation · Robotics

OpenEQA

Free while signed in. Answers cite the passages they came from.

First page
OpenEQA
The curator’s take

Meta's OpenEQA is an open-vocabulary benchmark for embodied question answering: 1,600+ human-written questions across 180+ real-world environments, with a calibrated LLM-as-judge metric that tracks human agreement closely.

Key points
01

Benchmark setup: Questions cover episodic-memory use cases (smart glasses) and active-exploration use cases (mobile robots), demanding that agents reason about the environment they occupy rather than just a single image.

02

LLM-as-judge scoring: The paper introduces an automatic LLM-powered evaluation protocol with strong correlation to human judgment, solving the open-vocabulary scoring problem that blocks most EQA benchmarks from scaling.

03

Frontier-model performance: GPT-4V, Claude 3, and Gemini Pro significantly outperform text-only baselines, but their gains come mostly from object recognition - for several question categories, they barely beat a blind-LLM baseline.

04

Gap to humans: Across the board, state-of-the-art multimodal LLMs lag well behind human performance on OpenEQA, positioning the benchmark as a concrete target for embodied-agent research.

Every Monday
Get next week’s papers.

The same picks and the same summaries, in your inbox. Free, and 176 issues deep.

Subscribe on Substack