🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Oct 9, 2026
Evaluation · Agents

Coding-Agent Benchmarks Should Match Their Users' Task Flows

First page
Coding-Agent Benchmarks Should Match Their Users' Task Flows
The curator’s take

Igor Slinko, Yaroslav Golubev and Sergey Titov (JetBrains Research) compare 4,782 real coding-agent sessions from JetBrains IDEs with issue-derived benchmarks and propose SWE-TaskFlow to reshape benchmarks toward a measured interaction pattern.

Ask this paper

Key points
01

Production data. 33% of sessions have three or more user messages. In SWE-Bench Pro, new features and bug fixes make up 93.8% of tasks; in long production sessions those two intents cover well under half of user messages.

02

Task Flow. Users ask about project code, plan, review, refactor and run things, and switch between these types within a session. Three public interaction corpora show different Task Flows from each other, so no single distribution is realistic for everyone.

03

SWE-TaskFlow. Keeps the verified tasks and tests of an issue-derived benchmark but splits the prompt and adds verifiable repository QA to steer the interaction toward a target Task Flow, scored by a TaskFlow Alignment Score.

04

Pilot. On 700 SWE-Bench Pro tasks, solving sequentially in several steps roughly doubles agent cost without a stable change in resolve rate.

05

Recommendation. Benchmarks should name a target use case and calibrate to measurements from it, since the interaction protocol is itself an evaluation dimension.

Abstract

The evaluation of coding agents generally strives to be as realistic as possible. In our study, we collect 4,782 agent sessions of real software engineers in JetBrains IDEs, which we call Production Sessions. Since our subject is interactive agents, we study the sessions with at least three user messages (33% of the sample). These long sessions differ from issue-derived benchmark tasks in two ways: (i) user requests span a far wider mix of task types - questions about the project's code, planning, review, refactoring, execution - and (ii) users switch between types throughout a session. Long-session samples from three public interaction corpora exhibit markedly different Task Flows (the distributions of session lengths, task types, and type-to-type transitions), so no single interaction distribution is universally realistic: benchmarks should name a target use case and calibrate to measurements from it. We present SWE-TaskFlow, an approach for transforming any issue-derived benchmark: it preserves the verified tasks and tests while steering the interaction toward a target Task Flow through prompt splitting and verifiable repository QA, with a TaskFlow Alignment Score (TFAS) for selecting among generated trajectories. In a pilot on 700 SWE-Bench Pro tasks, solving the task sequentially in several steps approximately doubles agent cost without a stable change in resolve rate: the interaction protocol itself is an important dimension of evaluation.

Every Monday
Get next week’s papers.
Subscribe on Substack