🚀NEW LABGetting Started with Claude AgentsStart lab
Reinforcement Learning · Robotics · Data

A Pinch of Human Data

First page
A Pinch of Human Data
Paper summary

Self-play reinforcement learning can train driving policies with no human data at all, swapping expensive human demonstrations for cheap large-scale simulation. The catch is that pure self-play tends to discover effective but alien driving conventions that real people cannot work with, and the usual fixes lean on brittle reward engineering and domain randomization.

Ask this paper

Key points
01

Human data as a regularizer: Instead of discarding demonstrations or imitating them wholesale, the method treats human data as a regularization objective layered on top of a minimal safe goal-reaching reward, keeping behavior compatible with people without hand-tuning conventions.

02

A little goes a long way: Just 30 minutes of human demonstrations, roughly 2500 times fewer than comparable imitation learning approaches, is enough to pull self-play policies into human-compatible behavior.

03

Cheap to train: The resulting policies coordinate with held-out human trajectories and finish training in 15 hours on a single consumer-grade GPU, which keeps the recipe accessible rather than a frontier-lab luxury.

04

Why it matters: Behavioral alignment with humans is the hard part of deploying autonomous policies in shared environments, and this work shows that a tiny, well-placed dose of human data can fix what massive reward engineering struggles to, pointing to a cleaner path for human-AI coordination.

Every Monday
Get next week’s papers.
Subscribe on Substack