🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
Training · Safety

Poisoning Language Models During Instruction Tuning

Free while signed in. Answers cite the passages they came from.

First page
Poisoning Language Models During Instruction Tuning
The curator’s take

Shows adversaries can poison LLMs via instruction tuning data.

Key points
01

Poisoning attack: Demonstrates adversaries can contribute poisoned examples to instruction tuning datasets to induce specific misbehaviors.

02

Cross-task poisoning: Poisoning can induce degenerate outputs across held-out tasks, not just the poisoned task - broad attack surface.

03

Supply-chain vulnerability: Highlights the supply-chain vulnerability of using community-sourced instruction data.

04

Alignment safety: Important for the field's thinking on data provenance and vetting for alignment datasets.

Every Monday
Get next week’s papers.

The same picks and the same summaries, in your inbox. Free, and 176 issues deep.

Subscribe on Substack