🚀NEW LABGetting Started with Claude AgentsStart lab
Evaluation · Multimodal · Safety

Red Teaming Visual Language Models

First page
Red Teaming Visual Language Models
Paper summary

Introduces the first dedicated red-teaming benchmark for VLMs, covering vulnerabilities unique to multimodal inputs.

Ask this paper

Key points
01

10-subtask benchmark: Probes vulnerabilities like image-based misdirection, multimodal jailbreaks, face fairness, and privacy leakage - a broader attack surface than text-only red teaming.

02

Open VLMs lag GPT-4V: 10 prominent open-source VLMs show meaningful weaknesses across the suite, with up to a 31% performance gap against GPT-4V on the red-teaming axis.

03

Red-teaming SFT works: Applying supervised fine-tuning on the proposed red-teaming dataset to LLaVA-v1.5 lifts test-set performance by ~10% without harming standard capabilities.

04

Practical alignment recipe: Demonstrates that targeted red-teaming data collection + SFT is a low-cost way to harden open VLMs against known multimodal attack vectors.

Every Monday
Get next week’s papers.
Subscribe on Substack