🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
Evaluation · Safety

MACHIAVELLI Benchmark

Free while signed in. Answers cite the passages they came from.

First page
MACHIAVELLI Benchmark
The curator’s take

A benchmark of 134 text-based Choose-Your-Own-Adventure games for measuring ethical trade-offs.

Key points
01

134 interactive games: Uses 134 text adventures with ~500K scenarios to evaluate agent behavior in rich social/ethical contexts.

02

Reward vs. ethics trade-off: Specifically measures how agents trade off goal-achievement (rewards) against ethical behavior (harm, deception, power-seeking).

03

Dark side measurement: Surfaces unethical behaviors like deception, manipulation, and power-seeking that may emerge when agents optimize for rewards.

04

Agent safety research: A foundational benchmark for the emerging "agent safety" sub-field in 2023-2024.

Every Monday
Get next week’s papers.

The same picks and the same summaries, in your inbox. Free, and 176 issues deep.

Subscribe on Substack