🚀NEW LABGetting Started with Claude AgentsStart lab
Training · Reinforcement Learning · Reasoning

The Era of Real-World Human Interaction

First page
The Era of Real-World Human Interaction
Paper summary

This work presents a post-training recipe that learns directly from real user conversations instead of static annotator labels. RLHI combines user-guided rewrites (using follow-ups as corrections) with persona-based rewards (ranking sampled candidates via a persona-conditioned reward model). Trained on WildChat conversations, it shows strong improvements in personalization, instruction following, and even transfers to reasoning tasks.

Ask this paper

Key points
01

Personas are distilled from long-term user histories and prepended at inference; training uses persona-conditioned DPO on rewrites and reward-ranked pairs.

02

Real chats contain rich correction signals, especially in later turns, providing dense supervision.

03

On WildChat-based evaluation, rewrites improve personalization and preference, while persona-based rewards lead in instruction following.

04

Benchmarks show strong results: 77.9% win rate on AlpacaEval 2.0, competitive on Arena-Hard, and reasoning accuracy rising from 26.5 to 31.8 across math/science datasets.

05

Key ablations: RL > SFT for interaction data, strong quality filters are essential, and user diversity matters more than depth per user.

06

Next steps include online continual learning, safer reward modeling, and privacy-preserving personalization.

Every Monday
Get next week’s papers.
Subscribe on Substack