Schrödinger's Code Repository: Have LLMs Learned SWE-bench or Memorized It?

Silin Chen, Xiaodong Gu and colleagues at Shanghai Jiao Tong University propose SchrodingerRepo, which rewrites the test repository at evaluation time so that coding agents cannot lean on memorized names, layouts and code patterns from SWE-bench's popular source repositories.
Ask this paper
Four transformation levels. problem-statement reconstruction, namespace remapping, intra-file layout reordering and functionality-preserving code rewriting. Executable behavior is preserved, so the same tests still decide success.
SWE-bench Verified drop. With all transformations on, Pass@1 falls by 6.0 to 14.4 points across GPT-5.4-mini, GPT-5.1, DeepSeek-v4-Flash and Gemini-3.1-Flash-Lite. Namespace remapping causes the largest single drop.
Cost moves more than accuracy. 81.6% to 83.6% of the extra actions go to exploration and localization, and token use more than doubles for several models under the full transformation.
Control that isolates memorization. On SWE-rebench instances created after the models' release, the same transformations leave Pass@1 unchanged while still raising interaction cost, which indicates the Verified drop comes from lost familiarity rather than harder tasks.
Beyond issue fixing. On SWE-QA, transformed repositories lower answer quality by up to 4.64 points and raise actions by 18.15% to 43.02%.
Abstract
Repository-level coding benchmarks have become the standard for evaluating coding agents, yet they inherently suffer from data leakage because they are built upon popular open-source repositories repeatedly used for training. Consequently, strong performance may reflect memorization of canonical repository cues rather than robust repository reasoning. We propose SchrodingerRepo (Schrödinger's Repository), an evaluation framework for testing coding agents under dynamically instantiated repository representations. Instead of repeatedly using a static representation of the test repository, SchrodingerRepo treats the test repository as an evaluation-time latent variable that is dynamically instantiated only when the agent enters the evaluation environment. The instantiated repository preserves the original executable behavior while eroding familiar cues such as naming conventions, file layouts, and implementation patterns through four transformation levels: problem statement reconstruction, namespace remapping, intra-file layout reordering, and functionality-preserving code rewriting. We evaluate popular LLMs on SWE-bench Verified and SWE-QA. Results show that removing familiar repository cues consistently degrades agent performance and substantially increases interaction costs across models. Further analysis reveals that the additional cost is primarily caused by increased difficulty in repository exploration and localization. These findings suggest that current coding agents may partially rely on memorized repository-side cues, highlighting the need for evaluation under dynamically instantiated repository representations.