🚀NEW LABGetting Started with Claude AgentsStart lab
Data · Safety · Training

Synthetic Data Reduces Sycophancy

First page
Synthetic Data Reduces Sycophancy
Paper summary

Google shows that fine-tuning on simple synthetic data can significantly reduce LLM sycophancy.

Ask this paper

Key points
01

Sycophancy problem: Sycophancy occurs when LLMs align their responses with perceived user views even when those views are factually incorrect.

02

Synthetic anti-sycophancy data: Constructs simple synthetic examples where the correct answer contradicts the user's stated view, then fine-tunes models on them.

03

Meaningful reduction: Fine-tuning on this synthetic data measurably reduces sycophantic behavior without degrading overall helpfulness.

04

Broader lesson: Offers a cheap, targeted intervention for a specific alignment failure mode - a template for addressing other narrow failure modes through targeted synthetic data.

Every Monday
Get next week’s papers.
Subscribe on Substack