🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
Evaluation · Agents · Multimodal

DataSpace

Free while signed in. Answers cite the passages they came from.

First page
DataSpace
The curator’s take

Real organizational analytics scatters evidence across databases, structured files, long documents, and video. Existing benchmarks isolate structured querying, retrieval, or open-ended analysis, leaving heterogeneous evidence discovery, complete tabular outputs, and deterministic scoring untested together.

Key points
01

Workspace-scale tasks: DataSpace contains 410 cross-language tasks over 7,439 artifacts totaling 15.01 GB across CSV, JSON, SQLite, Markdown, PDF, and video, and each agent receives only a question and a workspace before returning the full requested tabular result.

02

Deterministic evaluation: Scoring performs header-invariant column alignment, type-aware and precision-aware normalization, and order-aware row comparison, which removes the judge model from the loop entirely.

03

Harness choice is worth 15 points: Across six recently released frontier multimodal models and five widely used agent harnesses, the best accuracy reaches 66.34%, and holding the backbone fixed while swapping the harness moves accuracy by 15.36 points.

04

Why it matters: Multimodal evidence integration and joins reduce accuracy across all six backbones, so the benchmark remains far from saturated. It also served as the official evaluation benchmark for the KDD Cup 2026 Data Agents competition.

Every Monday
Get next week’s papers.

The same picks and the same summaries, in your inbox. Free, and 176 issues deep.

Subscribe on Substack