Red Teaming Visual Language Models
Free while signed in. Answers cite the passages they came from.

Introduces the first dedicated red-teaming benchmark for VLMs, covering vulnerabilities unique to multimodal inputs.
10-subtask benchmark: Probes vulnerabilities like image-based misdirection, multimodal jailbreaks, face fairness, and privacy leakage - a broader attack surface than text-only red teaming.
Open VLMs lag GPT-4V: 10 prominent open-source VLMs show meaningful weaknesses across the suite, with up to a 31% performance gap against GPT-4V on the red-teaming axis.
Red-teaming SFT works: Applying supervised fine-tuning on the proposed red-teaming dataset to LLaVA-v1.5 lifts test-set performance by ~10% without harming standard capabilities.
Practical alignment recipe: Demonstrates that targeted red-teaming data collection + SFT is a low-cost way to harden open VLMs against known multimodal attack vectors.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack