Design2Code
Free while signed in. Answers cite the passages they came from.

Design2Code tackles the front-end engineering problem of turning a visual design into working HTML/CSS and gives the community both a benchmark and strong MLLM baselines.
484-webpage benchmark: A curated set of 484 real-world webpages with screenshot + reference code pairs, paired with automatic metrics validated against human judgments.
Frontier MLLM comparison: GPT-4V comes out on top, Gemini Pro Vision is competitive, and an open-source fine-tuned Design2Code model matches Gemini Pro Vision - demonstrating that open models can be closed the gap on this task.
Failure modes: Models mostly fail on recalling fine-grained visual elements and reproducing correct layouts, not on individual snippets, isolating the remaining hard problem.
Prompting matters: The authors develop multimodal prompting strategies that materially lift GPT-4V and Gemini performance, giving practitioners ready-made recipes.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack