🚀NEW LABGetting Started with Claude AgentsStart lab
Reinforcement Learning · Multimodal · Safety

LLaVA-RLHF

First page
LLaVA-RLHF
Paper summary

Adapts factually augmented RLHF to aligning large multimodal models, reducing hallucination without falling into reward-hacking pitfalls.

Ask this paper

Key points
01

Factually augmented RLHF: Augments the reward model with factual-consistency signals (e.g., grounded-in-image checks), reducing the reward hacking common in vanilla multimodal RLHF.

02

Hallucination reduction: Produces meaningful reductions in hallucination on multimodal benchmarks compared to SFT-only or vanilla RLHF variants.

03

94% of text GPT-4: Reaches 94% of the performance level of text-only GPT-4 on LLaVA-Bench - closing a substantial gap via alignment alone.

04

Open recipe: Releases the full training recipe so the multimodal RLHF approach can be applied to other open VLMs.

Every Monday
Get next week’s papers.
Subscribe on Substack