Fine-Grained RLHF
First page

Paper summary
Trains LMs with segment-level human feedback rather than whole-response preferences.
Ask this paper
01
Segment-level rewards: Provides multiple reward models targeting specific dimensions (factuality, relevance, fluency) at the span level.
02
Long-form QA gains: Substantial improvements on long-form question answering where whole-response preferences are too coarse.
03
Toxicity reduction: Enables targeted reduction of toxic spans without degrading overall response quality.
04
Controllable RLHF: Enables model customization by emphasizing different reward dimensions at inference time.