Model Evaluation for Extreme Risks
Free while signed in. Answers cite the passages they came from.

DeepMind's framework for evaluating models for catastrophic-risk capabilities.
Dangerous-capability evaluation: Argues for evaluations targeting specifically dangerous capabilities (cyberattacks, bioweapons, manipulation) rather than general performance.
Responsible decisions: Connects evaluation results to decisions about training, deployment, access control, and security investments.
Red-team integration: Builds on dangerous-capability red-teaming methodology, formalizing it for frontier model governance.
Governance influence: Directly informed the UK AI Safety Institute's frontier model evaluation framework and similar efforts.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack