🚀NEW LABGetting Started with Claude AgentsStart lab
Agents · Safety

SocialRL

First page
SocialRL
Paper summary

The dispositions that make an assistant pleasant make it a poor delegate. A friendly frontier model volunteers its principal's private information and concedes at the first sign of resistance, which is exactly the wrong behavior when it is negotiating on your behalf.

Ask this paper

Key points
01

Trained where it matters: SocialRL trains social reasoning directly in a 4B model across six principal-driven domains including negotiation, job interviews, and marketplace haggling, with private information and asynchronous multi-agent interaction built into the environment.

02

The behavioral shift is stark: After training, 78% of buyer openings anchor below target, against 3% untrained. That is a learned strategic prior rather than a prompting trick.

03

Small model, better outcome: Cascade RL and multi-teacher distillation consolidate the specialists into a single 4B model at 0.627 average utility, above GPT-5.1 at 0.619 and GPT-5.2 at 0.613.

04

Why it matters: Assistant alignment and delegate alignment are not the same objective, and this is the first paper to make that gap measurable. As agents start transacting on behalf of users, the friendly-by-default posture stops being a safety feature and starts being a leak.

Every Monday
Get next week’s papers.
Subscribe on Substack