🚀NEW LABGetting Started with Claude AgentsStart lab
Safety

Math Jailbreaking Prompts

First page
Math Jailbreaking Prompts
Paper summary

uses GPT-4o to generate mathematically encoded prompts that serve as an effective jailbreaking technique; shows an average attack success rate of 73.6% across 13 state-of-the-art; this highlights the inability of existing safety training mechanisms to generalize to mathematically encoded inputs.

Ask this paper

Every Monday
Get next week’s papers.
Subscribe on Substack