🚀NEW LABGetting Started with Claude AgentsStart lab
Agents

LLM Agents Can Autonomously Hack Websites

First page
LLM Agents Can Autonomously Hack Websites
Paper summary

The paper shows GPT-4 agents with tool use and long context can autonomously exploit real websites, including performing blind SQL injection and schema extraction.

Ask this paper

Key points
01

Attack capabilities: GPT-4 agents autonomously perform SQL injection, extract database schemas blindly, and chain together multi-step exploits without any prior knowledge of the specific vulnerability.

02

Frontier-only: Only GPT-4 demonstrates this capability; tested open-source models and GPT-3.5 fail, suggesting a clear capability threshold enabled by frontier scale.

03

Real-world vulnerabilities: GPT-4 successfully discovered vulnerabilities in real websites in the wild, not just in lab-constructed targets.

04

Safety implications: Provides concrete evidence that frontier LLM capabilities can directly translate into offensive cyber capability, informing open-release and deployment decisions.

Every Monday
Get next week’s papers.
Subscribe on Substack