🚀NEW LABGetting Started with Claude AgentsStart lab
Reasoning

Trading Test-Time Compute for Adversarial Robustness

First page
Trading Test-Time Compute for Adversarial Robustness
Paper summary

Shows preliminary evidence that giving reasoning models like o1-preview and o1-mini more time to "think" during inference can improve their defense against adversarial attacks. Experiments covered various tasks, from basic math problems to image classification, showing that increasing inference-time compute often reduces the success rate of attacks to near zero. The approach doesn't work uniformly across all scenarios, particularly with certain StrongREJECT benchmark tests, and controlling how models use their compute time remains challenging. Despite these constraints, the findings suggest a promising direction for improving AI security without relying on traditional adversarial training methods.

Ask this paper

Every Monday
Get next week’s papers.
Subscribe on Substack