MobileCybench: Evaluating Agent Vulnerability Discovery via Executable Probes

Andy K. Zhang and colleagues at Stanford and UC Berkeley (with Percy Liang, Dan Boneh, Dawn Song and Ion Stoica) introduce MobileCybench, a benchmark that scores agent-reported exploits by replaying them and running executable probes that check whether a specific security property was violated.
Ask this paper
Probes check properties. Each probe encodes a security property of the application instead of a known bug, so it can confirm vulnerabilities nobody knew about when the probe was written. The benchmark has 495 author-written probes across 13 Android applications.
Four attack settings. Agents act as a malicious app on the victim's device or as a remote low-privilege attacker, with either an obfuscated APK or the source code.
Results. With only the obfuscated APK, OpenCode with GPT-5.6-Sol triggers probes in 53.8% of applications as a malicious app and 16.7% as a remote attacker. Source access raises the overall trigger rate only from 28.8% to 32.8%.
Refusals. Claude Code refused in 25.0% of runs with Opus 4.8 and 42.4% with Opus 5, despite cyber-verified access. Every trigger came from an application-specific probe; the generic probes never fired.
Real bugs found. Building and running the benchmark surfaced 23 previously unreported vulnerabilities, most confirmed by maintainers.
Abstract
AI agents now report vulnerabilities faster than maintainers can review them. Reports often depend on security properties specific to the application, and require considerable human labor to process. To mitigate this, we introduce a framework for evaluating vulnerability reports via probes, executable checks of security properties. A reported exploit is evaluated by replaying it against the application and running the probes: a triggered probe indicates both that the exploit succeeded and which security property it violated. As a probe encodes a security property rather than a known vulnerability, it can detect vulnerabilities that were not known when the probe was written. We instantiate the framework as MobileCybench, a benchmark for vulnerability discovery by AI agents in 13 Android applications, with 495 probes written and reviewed by the authors. We evaluate 5 coding agents (OpenCode with GPT-5.5, GPT-5.6-Sol, and GLM-5.2; Claude Code with Opus 4.8 and Opus 5) under 4 settings: as a malicious app on the victim's device or as a remote attacker with a low-privilege account, each with either only an obfuscated APK or access to the application's source code. Given only the obfuscated APK, the top agent, OpenCode with GPT-5.6-Sol, triggers probes in 53.8% of applications in the malicious-app setting and 16.7% in the remote-attacker setting. With source code, the trigger rate across all agents and both attack settings increases from 28.8% to 32.8%. Building and running the benchmark surfaced 23 previously unreported vulnerabilities, the majority of which have been confirmed by maintainers.