Quantifying Prompt-Format Sensitivity

CMU researchers show that LLM few-shot performance is shockingly sensitive to superficial prompt-formatting choices.
Ask this paper
Up to 76-point swings: On a Llama 2 13B model, subtle reformatting (changing delimiters, whitespace, capitalization) produces accuracy differences of up to 76 percentage points on classification tasks.
Broad model coverage: The effect holds across multiple open-source LLMs and persists at larger scale, though somewhat diminished.
Spurious features: What the authors call "spurious features" - choices irrelevant to the semantics of the task - are driving performance differences that researchers often attribute to method improvements.
Evaluation hygiene: Argues that every few-shot evaluation should marginalize over plausible formats or risk misattributing random variation to real capability differences.