A Pinch of Human Data
Free while signed in. Answers cite the passages they came from.

Self-play reinforcement learning can train driving policies with no human data at all, swapping expensive human demonstrations for cheap large-scale simulation. The catch is that pure self-play tends to discover effective but alien driving conventions that real people cannot work with, and the usual fixes lean on brittle reward engineering and domain randomization.
Human data as a regularizer: Instead of discarding demonstrations or imitating them wholesale, the method treats human data as a regularization objective layered on top of a minimal safe goal-reaching reward, keeping behavior compatible with people without hand-tuning conventions.
A little goes a long way: Just 30 minutes of human demonstrations, roughly 2500 times fewer than comparable imitation learning approaches, is enough to pull self-play policies into human-compatible behavior.
Cheap to train: The resulting policies coordinate with held-out human trajectories and finish training in 15 hours on a single consumer-grade GPU, which keeps the recipe accessible rather than a frontier-lab luxury.
Why it matters: Behavioral alignment with humans is the hard part of deploying autonomous policies in shared environments, and this work shows that a tiny, well-placed dose of human data can fix what massive reward engineering struggles to, pointing to a cleaner path for human-AI coordination.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack