LoLBench: Evaluating Coding Agents with Long-Horizon Proposals on Large Software Systems

Yun Peng, Zihan Wu and colleagues at Fudan University and City University of Hong Kong introduce LoLBench, a benchmark that tests coding agents on the full path from a human-written enhancement proposal to an implementation in a large codebase.
Ask this paper
Scale. 100 tasks across 29 software systems in five domains and several languages; proposals average about 5,000 words, systems 2.4 million lines of code, and reference PRs about 5,500 changed lines.
Two capabilities. Agents must turn user intent and high-level design into a specification (perception) and then write correct code (implementation).
Results. Of 28 agents, the best resolves 14% of tasks with a 52.7% fail-to-pass rate.
Bottleneck. Incomplete code localization is the main failure. Giving agents reference-derived file trees plus API specifications raises resolved rates by 16 to 22 points (2.4 to 17x), with a maximum of 34%.
Abstract
Modern coding agents can deliver increasingly large repository-level changes, and recent benchmarks reflect this by emphasizing long-horizon tasks with large reference implementations. Many benchmarks evaluate coding agents' implementation capability to produce correct code edits from detailed specifications. However, practical modular development tasks also require the perception capability of grounding user intent and high-level design to derive a specification. We introduce LoLBench to evaluate both capabilities through the entire proposal-to-implementation process on large software systems. It is a multilingual benchmark of 100 tasks across 29 software systems in five domains. Each task provides a human-written enhancement proposal with user intent and high-level design. On average, proposals contain about 5,000 words, software systems contain 2.4 million source lines of code (LoC), and implementation pull requests (PRs) change approximately 5,500 LoC. Across 28 agents we evaluated, the best agent resolves only 14% of tasks and achieves a 52.7% Fail-to-Pass (F2P) pass rate. Failure analysis identifies incomplete code localization as a major bottleneck, while providing reference-derived file trees alongside API specifications improves resolved rates by 16--22 percentage points (2.4--17$\times$), reaching at most 34%. These results show that both perception and implementation remain central challenges for coding agents in practical modular development on large software systems. LoLBench is available at https://huggingface.co/datasets/lolbench26/LoLBench.