FLASK
Free while signed in. Answers cite the passages they came from.

Proposes fine-grained evaluation of LLMs decomposed into 12 alignment skill sets.
12-skill taxonomy: Decomposes holistic LLM evaluation into skills like logical reasoning, factuality, commonsense, readability, harmlessness, etc.
Instance-level annotation: Each evaluation instance is labeled with which skills, domains, and difficulty levels it exercises, enabling fine-grained performance analysis.
Skill-specific insights: Reveals that models excel differently on different skills - useful for targeted model selection and iteration.
Evaluation paradigm shift: Part of the broader move from single-number benchmarks to multi-dimensional skill-based evaluation that shaped 2024's eval ecosystem.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack