Red Teaming Visual Language Models

Introduces the first dedicated red-teaming benchmark for VLMs, covering vulnerabilities unique to multimodal inputs.
Ask this paper
10-subtask benchmark: Probes vulnerabilities like image-based misdirection, multimodal jailbreaks, face fairness, and privacy leakage - a broader attack surface than text-only red teaming.
Open VLMs lag GPT-4V: 10 prominent open-source VLMs show meaningful weaknesses across the suite, with up to a 31% performance gap against GPT-4V on the red-teaming axis.
Red-teaming SFT works: Applying supervised fine-tuning on the proposed red-teaming dataset to LLaVA-v1.5 lifts test-set performance by ~10% without harming standard capabilities.
Practical alignment recipe: Demonstrates that targeted red-teaming data collection + SFT is a low-cost way to harden open VLMs against known multimodal attack vectors.