Poisoning Language Models During Instruction Tuning
Free while signed in. Answers cite the passages they came from.
First page

The curator’s take
Key pointsShows adversaries can poison LLMs via instruction tuning data.
01
Poisoning attack: Demonstrates adversaries can contribute poisoned examples to instruction tuning datasets to induce specific misbehaviors.
02
Cross-task poisoning: Poisoning can induce degenerate outputs across held-out tasks, not just the poisoned task - broad attack surface.
03
Supply-chain vulnerability: Highlights the supply-chain vulnerability of using community-sourced instruction data.
04
Alignment safety: Important for the field's thinking on data provenance and vetting for alignment datasets.
Every Monday
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack