🚀NEW LABGetting Started with Claude AgentsStart lab
Training · Safety

Poisoning Language Models During Instruction Tuning

First page
Poisoning Language Models During Instruction Tuning
Paper summary

Shows adversaries can poison LLMs via instruction tuning data.

Ask this paper

Key points
01

Poisoning attack: Demonstrates adversaries can contribute poisoned examples to instruction tuning datasets to induce specific misbehaviors.

02

Cross-task poisoning: Poisoning can induce degenerate outputs across held-out tasks, not just the poisoned task - broad attack surface.

03

Supply-chain vulnerability: Highlights the supply-chain vulnerability of using community-sourced instruction data.

04

Alignment safety: Important for the field's thinking on data provenance and vetting for alignment datasets.

Every Monday
Get next week’s papers.
Subscribe on Substack