Quantifying Prompt-Format Sensitivity
Free while signed in. Answers cite the passages they came from.

CMU researchers show that LLM few-shot performance is shockingly sensitive to superficial prompt-formatting choices.
Up to 76-point swings: On a Llama 2 13B model, subtle reformatting (changing delimiters, whitespace, capitalization) produces accuracy differences of up to 76 percentage points on classification tasks.
Broad model coverage: The effect holds across multiple open-source LLMs and persists at larger scale, though somewhat diminished.
Spurious features: What the authors call "spurious features" - choices irrelevant to the semantics of the task - are driving performance differences that researchers often attribute to method improvements.
Evaluation hygiene: Argues that every few-shot evaluation should marginalize over plausible formats or risk misattributing random variation to real capability differences.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack