Synthetic Data Reduces Sycophancy
First page

Paper summary
Google shows that fine-tuning on simple synthetic data can significantly reduce LLM sycophancy.
Ask this paper
01
Sycophancy problem: Sycophancy occurs when LLMs align their responses with perceived user views even when those views are factually incorrect.
02
Synthetic anti-sycophancy data: Constructs simple synthetic examples where the correct answer contradicts the user's stated view, then fine-tunes models on them.
03
Meaningful reduction: Fine-tuning on this synthetic data measurably reduces sycophantic behavior without degrading overall helpfulness.
04
Broader lesson: Offers a cheap, targeted intervention for a specific alignment failure mode - a template for addressing other narrow failure modes through targeted synthetic data.