🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Sep 19, 2026
Agents · Code

The Missing Complement: State-Conditioned Minimal Sufficient Evidence for Coding Agents

First page
The Missing Complement: State-Conditioned Minimal Sufficient Evidence for Coding Agents
The curator’s take

Zhexi Feng, Ruiyi Zhang, Yongbo Yang and Pengtao Xie formulate state-conditioned minimal sufficient evidence recovery, where the task is to supply the support an agent's next decision still lacks given what it has already read, and build SERBench and MSS-Complement.

Ask this paper

Key points
01

Sufficiency is a property of the set. A ranker can fill its budget with variants of one required fact and leave the decision unsupported, which per-passage relevance scoring cannot detect.

02

SERBench scores 500 held-out agent states. Drawn from 45 repositories, recording what the agent has seen and crediting only sets that cover every fact the annotated decision requires.

03

Set construction beats ranking by 11.6 points at five items. MSS-Complement recovers a complete set for 73.0 percent of states at five items and 80.6 percent at eight, against 61.4 and 72.4 percent for Qwen3 embedding with reranking; a similarity-only control reaches 66.6 percent, placing the gain in the set-level policy rather than the underlying computation.

04

It answers from a 76.2 percent smaller prompt. On AMA-Bench, with accuracy 2.08 points above that benchmark's own memory agent.

05

Missing one group is expensive. Removing a single required group from an otherwise complete set costs 12.3 and 11.1 points of repair-localization precision under two executors.

Abstract

A coding agent halfway through an issue has already read much of what a retriever ranks highest. Relevance is scored per passage, but sufficiency belongs to the set: a ranker can fill its budget with variants of one required fact and leave the decision unsupported. We formulate state-conditioned minimal sufficient evidence recovery: given a captured agent state, recover a compact evidence combination that supplies the support its next decision still lacks. SERBench measures this on 500 held-out states from 45 repositories, recording what the agent has seen and crediting only sets that cover every fact the current decision was annotated to require. MSS-Complement treats acquisition as set construction, not ranking. Three semantic calls propose a jointly sufficient set, search for what it lacks, and return 4-8 intact source units within 6,144 tokens. One configuration, fixed on calibration data, recovers a complete set for 73.0% of those states at five items and 80.6% at eight, against 61.4% and 72.4% for Qwen3 embedding with reranking. A matched control ranking by similarity alone reaches 66.6%, placing the gain in the set-level policy, not the computation. From frozen repository source with no gold-derived pool, the lead is 5.0 points. On AMA-Bench it answers from a 76.2% smaller answer prompt, with accuracy 2.08 points above that benchmark's own memory agent. Removing one required group from an otherwise complete set costs 12.3 and 11.1 points of repair-localization precision under two executors. Retrieval for agents is better posed as recovering what a decision lacks than re-ranking what an issue resembles.

Every Monday
Get next week’s papers.
Subscribe on Substack