🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
Evaluation

Quantifying Prompt-Format Sensitivity

Free while signed in. Answers cite the passages they came from.

First page
Quantifying Prompt-Format Sensitivity
The curator’s take

CMU researchers show that LLM few-shot performance is shockingly sensitive to superficial prompt-formatting choices.

Key points
01

Up to 76-point swings: On a Llama 2 13B model, subtle reformatting (changing delimiters, whitespace, capitalization) produces accuracy differences of up to 76 percentage points on classification tasks.

02

Broad model coverage: The effect holds across multiple open-source LLMs and persists at larger scale, though somewhat diminished.

03

Spurious features: What the authors call "spurious features" - choices irrelevant to the semantics of the task - are driving performance differences that researchers often attribute to method improvements.

04

Evaluation hygiene: Argues that every few-shot evaluation should marginalize over plausible formats or risk misattributing random variation to real capability differences.

Every Monday
Get next week’s papers.

The same picks and the same summaries, in your inbox. Free, and 176 issues deep.

Subscribe on Substack