🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
← All papers  /  Sep 3, 2026
Code · Agents · Evaluation

SWE-Gate: Passing Functional Tests Is Not Enough for Software Engineering Agents

First page
SWE-Gate: Passing Functional Tests Is Not Enough for Software Engineering Agents
The curator’s take

Xin He and colleagues at Sun Yat-sen University introduce SWE-Gate, a repository-level benchmark that scores coding agents on review-derived acceptance constraints alongside functional tests, and shows that passing the tests is far from passing review.

Ask this paper

Key points
01

221 of 644 functionally passing repairs fail review constraints: the headline number. Functional-only evaluation materially overestimates what an agent has actually delivered.

02

Constraints mined from real PR review comments: 303 repository-level repair instances across 75 open-source Python repos, each with separate functional and constraint tests plus non-compliant and gold patches.

03

Clean separation of two abilities: the design lets you measure issue-resolution capability independently from compliance with the full repair specification, which SWE-bench-style benchmarks conflate.

04

Four backends under one scaffold: results hold across capability levels, so the gap is not an artifact of a weak model.

05

Why it matters: the practical bottleneck on shipping agent patches is review, not tests. This is the first benchmark that scores the thing that actually blocks the merge.

Abstract

Repository-level software engineering benchmarks have significantly advanced the evaluation of coding agents, but existing benchmarks primarily measure whether generated patches pass functional tests and overlook review-derived acceptance constraints (review constraints) that often influence whether a patch is acceptable in real-world software development. We introduce SWE-Gate, a repository-level benchmark for software engineering agents that explicitly evaluates review constraint compliance alongside functional correctness. SWE-Gate derives review constraints from real pull request review comments and synthesizes repository-level repair instances around these constraints. Each instance provides separate functional and constraint tests, together with non-compliant and gold patches, enabling explicit separation between issue resolution capability and review constraint compliance. We construct SWE-Gate with 303 repository-level repair instances spanning 75 open-source Python repositories across diverse software domains. Experiments with four LLM backends spanning different capability levels under a common coding-agent scaffold reveal a substantial gap between functional success and success under the complete repair specification: among 644 repairs that pass the functional tests, 221 fail to satisfy the provided review constraints. These findings show that functional-only evaluation overestimates agents' ability to satisfy the full requirements of repository-level repair tasks. The replication package including code, data, and experimental results is available at https://github.com/DeepSoftwareAnalytics/SWE-Gate.

Every Monday
Get next week’s papers.
Subscribe on Substack