Poisoning Language Models During Instruction Tuning
First page

Paper summary
Shows adversaries can poison LLMs via instruction tuning data.
Ask this paper
01
Poisoning attack: Demonstrates adversaries can contribute poisoned examples to instruction tuning datasets to induce specific misbehaviors.
02
Cross-task poisoning: Poisoning can induce degenerate outputs across held-out tasks, not just the poisoned task - broad attack surface.
03
Supply-chain vulnerability: Highlights the supply-chain vulnerability of using community-sourced instruction data.
04
Alignment safety: Important for the field's thinking on data provenance and vetting for alignment datasets.