🚀NEW LABGetting Started with Claude AgentsStart lab
Safety

Leaky Thoughts

First page
Leaky Thoughts
Paper summary

This work explores how reasoning traces in large reasoning models (LRMs) leak private user data, despite being assumed internal and safe. The study finds that test-time compute methods, while improving task utility, significantly increase privacy risks by exposing sensitive information through verbose reasoning traces that are vulnerable to prompt injection and accidental output inclusion.

Ask this paper

Every Monday
Get next week’s papers.
Subscribe on Substack