🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
Training · Reinforcement Learning · Reasoning

The Era of Real-World Human Interaction

Free while signed in. Answers cite the passages they came from.

First page
The Era of Real-World Human Interaction
The curator’s take

This work presents a post-training recipe that learns directly from real user conversations instead of static annotator labels. RLHI combines user-guided rewrites (using follow-ups as corrections) with persona-based rewards (ranking sampled candidates via a persona-conditioned reward model). Trained on WildChat conversations, it shows strong improvements in personalization, instruction following, and even transfers to reasoning tasks.

Key points
01

Personas are distilled from long-term user histories and prepended at inference; training uses persona-conditioned DPO on rewrites and reward-ranked pairs.

02

Real chats contain rich correction signals, especially in later turns, providing dense supervision.

03

On WildChat-based evaluation, rewrites improve personalization and preference, while persona-based rewards lead in instruction following.

04

Benchmarks show strong results: 77.9% win rate on AlpacaEval 2.0, competitive on Arena-Hard, and reasoning accuracy rising from 26.5 to 31.8 across math/science datasets.

05

Key ablations: RL > SFT for interaction data, strong quality filters are essential, and user diversity matters more than depth per user.

06

Next steps include online continual learning, safer reward modeling, and privacy-preserving personalization.

Every Monday
Get next week’s papers.

The same picks and the same summaries, in your inbox. Free, and 176 issues deep.

Subscribe on Substack