🚀NEW LABGetting Started with Claude AgentsStart lab
Evaluation

Quantifying Prompt-Format Sensitivity

First page
Quantifying Prompt-Format Sensitivity
Paper summary

CMU researchers show that LLM few-shot performance is shockingly sensitive to superficial prompt-formatting choices.

Ask this paper

Key points
01

Up to 76-point swings: On a Llama 2 13B model, subtle reformatting (changing delimiters, whitespace, capitalization) produces accuracy differences of up to 76 percentage points on classification tasks.

02

Broad model coverage: The effect holds across multiple open-source LLMs and persists at larger scale, though somewhat diminished.

03

Spurious features: What the authors call "spurious features" - choices irrelevant to the semantics of the task - are driving performance differences that researchers often attribute to method improvements.

04

Evaluation hygiene: Argues that every few-shot evaluation should marginalize over plausible formats or risk misattributing random variation to real capability differences.

Every Monday
Get next week’s papers.
Subscribe on Substack