🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
← All papers  /  Sep 10, 2026
Agents

Can AI Agents Detect and Repair Artifact Drift in Network Experiments?

First page
Can AI Agents Detect and Repair Artifact Drift in Network Experiments?
The curator’s take

Tianzhu Zhang (Nokia Bell Labs), Weichen Tao (Telecom Paris), Changgang Zheng (Nanjing University) and colleagues define artifact integrity as a property an agent must preserve, and build NetArtifactBench to measure whether agents can repair inconsistent experiment records without breaking supported claims.

Ask this paper

Key points
01

The property being introduced: An agent operating on a network experiment record should not be judged only on task completion. The record's claims must stay supported by available evidence, confined to the scope that evidence establishes, and traceable through the artifacts encoding their support.

02

The benchmark: 52 instances built from public network-system artifacts, with injected inconsistencies ranging from direct contradictions to unstated relations spread across several artifacts, scored deterministically.

03

The headline gap: Average contract pass rate is 65.3% across 5,980 outputs, but no agent runtime exceeds 30% when repair requires recovering implicit relations and propagating the change across artifacts.

04

What that boundary separates: Local correction, which agents handle, from complete record-level repair, which they do not.

05

Scope: 23 agent configurations across three general-purpose agent runtimes.

Abstract

In recent years, AI agents have evolved into capable assistants that carry out multi-step tasks in digital environments. The network systems community is beginning to explore these capabilities in operational and experimental settings. However, an agent operating in network systems should not be judged solely by whether it completes the immediate task. The experiment record it modifies must also remain trustworthy. We call this property artifact integrity: the record's claims must remain supported by the available evidence, confined to the scope established by that evidence, and traceable through the artifacts that encode their support. To make this property measurable, we introduce NetArtifactBench, which tests whether AI agents can repair inconsistent records derived from public network-system artifacts while preserving claims that remain supported. The benchmark contains 52 instances with injected inconsistencies ranging from direct contradictions to unstated relations spread across several artifacts. We evaluate 23 agent configurations across three general-purpose AI agent runtimes using deterministic scoring. The average contract pass rate is 65.3 % across 5,980 outputs, but no agent runtime exceeds 30 % when repair requires recovering implicit relations and propagating changes across artifacts. These results reveal a sharp boundary between local correction and complete record-level repair. Therefore, we argue that artifact integrity should become a first-class design and evaluation requirement for AI agents operating on network systems.

Every Monday
Get next week’s papers.
Subscribe on Substack