SocialRL
Free while signed in. Answers cite the passages they came from.

The dispositions that make an assistant pleasant make it a poor delegate. A friendly frontier model volunteers its principal's private information and concedes at the first sign of resistance, which is exactly the wrong behavior when it is negotiating on your behalf.
Trained where it matters: SocialRL trains social reasoning directly in a 4B model across six principal-driven domains including negotiation, job interviews, and marketplace haggling, with private information and asynchronous multi-agent interaction built into the environment.
The behavioral shift is stark: After training, 78% of buyer openings anchor below target, against 3% untrained. That is a learned strategic prior rather than a prompting trick.
Small model, better outcome: Cascade RL and multi-teacher distillation consolidate the specialists into a single 4B model at 0.627 average utility, above GPT-5.1 at 0.619 and GPT-5.2 at 0.613.
Why it matters: Assistant alignment and delegate alignment are not the same objective, and this is the first paper to make that gap measurable. As agents start transacting on behalf of users, the friendly-by-default posture stops being a safety feature and starts being a leak.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack