🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
Agents · Safety

SocialRL

Free while signed in. Answers cite the passages they came from.

First page
SocialRL
The curator’s take

The dispositions that make an assistant pleasant make it a poor delegate. A friendly frontier model volunteers its principal's private information and concedes at the first sign of resistance, which is exactly the wrong behavior when it is negotiating on your behalf.

Key points
01

Trained where it matters: SocialRL trains social reasoning directly in a 4B model across six principal-driven domains including negotiation, job interviews, and marketplace haggling, with private information and asynchronous multi-agent interaction built into the environment.

02

The behavioral shift is stark: After training, 78% of buyer openings anchor below target, against 3% untrained. That is a learned strategic prior rather than a prompting trick.

03

Small model, better outcome: Cascade RL and multi-teacher distillation consolidate the specialists into a single 4B model at 0.627 average utility, above GPT-5.1 at 0.619 and GPT-5.2 at 0.613.

04

Why it matters: Assistant alignment and delegate alignment are not the same objective, and this is the first paper to make that gap measurable. As agents start transacting on behalf of users, the friendly-by-default posture stops being a safety feature and starts being a leak.

Every Monday
Get next week’s papers.

The same picks and the same summaries, in your inbox. Free, and 176 issues deep.

Subscribe on Substack