FLASK
First page

Paper summary
Proposes fine-grained evaluation of LLMs decomposed into 12 alignment skill sets.
Ask this paper
01
12-skill taxonomy: Decomposes holistic LLM evaluation into skills like logical reasoning, factuality, commonsense, readability, harmlessness, etc.
02
Instance-level annotation: Each evaluation instance is labeled with which skills, domains, and difficulty levels it exercises, enabling fine-grained performance analysis.
03
Skill-specific insights: Reveals that models excel differently on different skills - useful for targeted model selection and iteration.
04
Evaluation paradigm shift: Part of the broader move from single-number benchmarks to multi-dimensional skill-based evaluation that shaped 2024's eval ecosystem.