🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Sep 25, 2026
Agents · Evaluation · Training

Trains but Doesn't Learn: A Post-Training Delivery Benchmark for LLM Agents as Forward-Deployed Engineers

First page
Trains but Doesn't Learn: A Post-Training Delivery Benchmark for LLM Agents as Forward-Deployed Engineers
The curator’s take

Weihang Ding (UC Berkeley) and Junfei Zhan (Imperial College London) build a benchmark in which an LLM agent acts as a forward-deployed engineer delivering fine-tuned models to customers, and measure whether it can be trusted to deliver rather than only raise a metric. Accepted to the EMNLP 2026 Industry Track.

Ask this paper

Key points
01

Delivery plane. The agent drives ten stages from customer ticket to deployment, under a budget, a human-approval gate and reproducibility requirements, and an oracle scores each stage from platform-recorded facts.

02

Trains but does not learn. The central silent failure is a run where loss falls and every check stays green but the delivered model is no better than the base.

03

Two defenses. An operator-run acceptance gate catches every such run before payment, and a detector calibrated on known-corrupted runs flags severe corruption mid-run using a forward pass on a 64-example probe every ten steps.

04

Frontier agents on real GPUs. Claude Opus 5, GPT-5.6-luna, Gemini 3.7 Flash and DeepSeek V4-Pro ran end to end on L40S, A100 and H200 GPUs with 8B to 70B bases; the flagship arms held the governance gate in 360 of 360 pressure trials and refused all 60 infeasible tickets.

05

Human comparison. Human engineer-plus-assistant arms had a lower judgment residual than every autonomous agent (0.06 and 0.09 against 0.10), and no arm reached zero.

Abstract

Post-training is becoming a service (PTaaS): a customer hands an operator data and a goal, and a forward-deployed engineer (FDE) returns a fine-tuned, evaluated, and deployed model under a budget, a human-approval gate, and reproducibility requirements. Seating an LLM agent in the FDE seat raises a question existing benchmarks cannot answer: not whether an agent can raise a metric, but whether it can be trusted to deliver. We answer it on a governed delivery plane, where an agent drives ten stages and an oracle scores each stage from platform-recorded facts. The central silent failure is the run that trains but does not learn (TBDL): loss falls, every signal stays green, and the delivered model is no better than the base. An operator-run acceptance gate catches every such run before payment, and a detector calibrated on known-corrupted runs flags severe corruption mid-run. We ran four frontier agents (Claude Opus 5, GPT-5.6-luna, Gemini 3.7 Flash, DeepSeek V4-Pro) end to end on metered L40S, A100, and H200 GPUs across 8B to 70B open bases, certifying every scenario before scoring. We also ran a human FDE arm under the same oracle and compare every agent against it.

Every Monday
Get next week’s papers.
Subscribe on Substack