🚀NEW LABGetting Started with Claude AgentsStart lab
Evaluation

Design2Code

First page
Design2Code
Paper summary

Design2Code tackles the front-end engineering problem of turning a visual design into working HTML/CSS and gives the community both a benchmark and strong MLLM baselines.

Ask this paper

Key points
01

484-webpage benchmark: A curated set of 484 real-world webpages with screenshot + reference code pairs, paired with automatic metrics validated against human judgments.

02

Frontier MLLM comparison: GPT-4V comes out on top, Gemini Pro Vision is competitive, and an open-source fine-tuned Design2Code model matches Gemini Pro Vision - demonstrating that open models can be closed the gap on this task.

03

Failure modes: Models mostly fail on recalling fine-grained visual elements and reproducing correct layouts, not on individual snippets, isolating the remaining hard problem.

04

Prompting matters: The authors develop multimodal prompting strategies that materially lift GPT-4V and Gemini performance, giving practitioners ready-made recipes.

Every Monday
Get next week’s papers.
Subscribe on Substack