🚀NEW LABGetting Started with Claude AgentsStart lab
Evaluation · Safety

FLASK

First page
FLASK
Paper summary

Proposes fine-grained evaluation of LLMs decomposed into 12 alignment skill sets.

Ask this paper

Key points
01

12-skill taxonomy: Decomposes holistic LLM evaluation into skills like logical reasoning, factuality, commonsense, readability, harmlessness, etc.

02

Instance-level annotation: Each evaluation instance is labeled with which skills, domains, and difficulty levels it exercises, enabling fine-grained performance analysis.

03

Skill-specific insights: Reveals that models excel differently on different skills - useful for targeted model selection and iteration.

04

Evaluation paradigm shift: Part of the broader move from single-number benchmarks to multi-dimensional skill-based evaluation that shaped 2024's eval ecosystem.

Every Monday
Get next week’s papers.
Subscribe on Substack