🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Oct 1, 2026
Agents · Code · Evaluation

LoLBench: Evaluating Coding Agents with Long-Horizon Proposals on Large Software Systems

First page
LoLBench: Evaluating Coding Agents with Long-Horizon Proposals on Large Software Systems
The curator’s take

Yun Peng, Zihan Wu and colleagues at Fudan University and City University of Hong Kong introduce LoLBench, a benchmark that tests coding agents on the full path from a human-written enhancement proposal to an implementation in a large codebase.

Ask this paper

Key points
01

Scale. 100 tasks across 29 software systems in five domains and several languages; proposals average about 5,000 words, systems 2.4 million lines of code, and reference PRs about 5,500 changed lines.

02

Two capabilities. Agents must turn user intent and high-level design into a specification (perception) and then write correct code (implementation).

03

Results. Of 28 agents, the best resolves 14% of tasks with a 52.7% fail-to-pass rate.

04

Bottleneck. Incomplete code localization is the main failure. Giving agents reference-derived file trees plus API specifications raises resolved rates by 16 to 22 points (2.4 to 17x), with a maximum of 34%.

Abstract

Modern coding agents can deliver increasingly large repository-level changes, and recent benchmarks reflect this by emphasizing long-horizon tasks with large reference implementations. Many benchmarks evaluate coding agents' implementation capability to produce correct code edits from detailed specifications. However, practical modular development tasks also require the perception capability of grounding user intent and high-level design to derive a specification. We introduce LoLBench to evaluate both capabilities through the entire proposal-to-implementation process on large software systems. It is a multilingual benchmark of 100 tasks across 29 software systems in five domains. Each task provides a human-written enhancement proposal with user intent and high-level design. On average, proposals contain about 5,000 words, software systems contain 2.4 million source lines of code (LoC), and implementation pull requests (PRs) change approximately 5,500 LoC. Across 28 agents we evaluated, the best agent resolves only 14% of tasks and achieves a 52.7% Fail-to-Pass (F2P) pass rate. Failure analysis identifies incomplete code localization as a major bottleneck, while providing reference-derived file trees alongside API specifications improves resolved rates by 16--22 percentage points (2.4--17$\times$), reaching at most 34%. These results show that both perception and implementation remain central challenges for coding agents in practical modular development on large software systems. LoLBench is available at https://huggingface.co/datasets/lolbench26/LoLBench.

Every Monday
Get next week’s papers.
Subscribe on Substack