LLM Agents Can Autonomously Hack Websites

The paper shows GPT-4 agents with tool use and long context can autonomously exploit real websites, including performing blind SQL injection and schema extraction.
Ask this paper
Attack capabilities: GPT-4 agents autonomously perform SQL injection, extract database schemas blindly, and chain together multi-step exploits without any prior knowledge of the specific vulnerability.
Frontier-only: Only GPT-4 demonstrates this capability; tested open-source models and GPT-3.5 fail, suggesting a clear capability threshold enabled by frontier scale.
Real-world vulnerabilities: GPT-4 successfully discovered vulnerabilities in real websites in the wild, not just in lab-constructed targets.
Safety implications: Provides concrete evidence that frontier LLM capabilities can directly translate into offensive cyber capability, informing open-release and deployment decisions.