🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Oct 10, 2026
Evaluation · Agents

BrickBench: Evaluating Agentic Brick Design

First page
BrickBench: Evaluating Agentic Brick Design
The curator’s take

Peter Kulits, Jiajun Wu and colleagues at Stanford (with Max Planck and Inria's Cordelia Schmid) introduce BrickBench, a benchmark where coding agents design LEGO assemblies from text prompts that must also be physically buildable.

Ask this paper

Key points
01

Benchmark. 300 prompts across three settings: Model (up to 400 parts), Set (400-4,000 parts) and Alt-Build (restricted to the part inventory of set 10698).

02

Metrics. Physical validity (connections, collisions, stability in simulation), semantic alignment via VQA over yes/no questions decomposed from each prompt, and design quality via VLM pairwise Elo validated against human raters.

03

BrickAgent. An environment for programmatic construction with part search, connector-based placement, subassemblies, rendering and a validator that names faulty parts. Without it, valid assemblies fall from 100% to 40% for GPT-6 Astra and below 1% for GPT-5.6 Luna.

04

Results. Eleven frontier agents and two data-driven baselines mostly satisfy physical and semantic requirements, but human raters pick out the human-designed assembly in 323 of 360 comparisons.

Abstract

We propose BrickBench, a benchmark for agentic text-conditioned LEGO-set design. Given a prompt, an agent is tasked with producing an assembly that not only satisfies semantic and design criteria, but that can also be physically built. To do so, it must select parts from a discrete library and reason jointly about local and global constraints. We score validity, alignment, and design across three settings that vary in scale and part availability. We provide BrickAgent, an environment for coding agents to construct, inspect, and validate their designs. We find that leading agents largely satisfy verifiable physical and semantic requirements, but fall short of human designs. We release our benchmark and environment at this http URL

Every Monday
Get next week’s papers.
Subscribe on Substack