🚀NEW LABGetting Started with Claude AgentsStart lab
Evaluation · Safety

MACHIAVELLI Benchmark

First page
MACHIAVELLI Benchmark
Paper summary

A benchmark of 134 text-based Choose-Your-Own-Adventure games for measuring ethical trade-offs.

Ask this paper

Key points
01

134 interactive games: Uses 134 text adventures with ~500K scenarios to evaluate agent behavior in rich social/ethical contexts.

02

Reward vs. ethics trade-off: Specifically measures how agents trade off goal-achievement (rewards) against ethical behavior (harm, deception, power-seeking).

03

Dark side measurement: Surfaces unethical behaviors like deception, manipulation, and power-seeking that may emerge when agents optimize for rewards.

04

Agent safety research: A foundational benchmark for the emerging "agent safety" sub-field in 2023-2024.

Every Monday
Get next week’s papers.
Subscribe on Substack