A Pinch of SFT, A Dash of RL: When Reinforcement Learning Helps Long-Horizon Advertising Agents

Aakash Kolekar, Sahika Genc and colleagues at Amazon Advertising and AWS Agentic AI study when to use SFT, RL or both for long-horizon advertising analytics agents, and turn the answer into a per-feature routing diagnostic (EMNLP 2026 Industry Track).
Ask this paper
Three regimes. Checkpoint trajectories fall into Imitation, where SFT captures reliable teacher behavior, Lift, where both stages help, and Discovery, where useful behavior lies outside teacher support.
Routing rule. Teacher support and reward-observable headroom decide whether a feature gets SFT only, SFT then RL, more RL, or more environment work; the rule predicted 15 of 18 later experiments.
Skill gains. On GPT-OSS 120B, targeted SFT then RL beat a frontier control on 7 of 8 advertiser skills; five gains had 95% intervals excluding zero and one skill regressed.
Leakage. Non-disclosure improved by 11.27 points, and targeted RL cut standard leakage from 11.8% to 2.9% and adversarial leakage from 22.9% to 6.8% with actionability roughly unchanged.
Compute. Against uniform RL with the same rewards, targeted RL raised the seven-skill mean delta from +1.62 to +3.57 with 43% less incremental RL compute.
Abstract
Enterprise analytics agents solve long-horizon tool-use problems over distributed business data, requiring retrieval, reasoning, API calls, code execution, and adaptation to intermediate observations. Supervised fine-tuning (SFT) calibrates tool syntax and teacher-supported behavior, whereas reinforcement learning (RL) can explore reward-supported behaviors beyond demonstrations; applied uniformly, however, RL can perturb already-calibrated skills. We study how to balance SFT and RL under production-mirroring beta APIs. We observe that, in our controlled experiment, checkpoint trajectories retrospectively separated into three regimes: Imitation, where SFT captured reliable teacher behavior; Lift, where both stages helped; and Discovery, where useful reward-observable behavior lay outside reliable teacher support. We leverage this prospectively, using teacher support and reward-observable headroom to route features to SFT only, SFT then RL, increased RL allocation, or further environment development. Across 18 subsequent feature-specific experiments, the diagnostic predicted 15/18 observed trajectories. On GPT-OSS 120B, targeted SFT then RL produced positive point estimates on 7/8 advertiser skills relative to a frontier Control; five positive gains had paired 95% confidence intervals excluding zero, while one skill had a confidence-supported regression. The largest gain was non-disclosure (+11.27 points; 95% CI [+9.72, +12.82]). A separate SME audit surfaced that targeted RL reduces standard leakage from 11.8% to 2.9% and adversarial leakage from 22.9% to 6.8% relative to SFT while preserving actionability (86.2% to 85.7%). In a matched uniform-versus-targeted comparison with shared rewards and optimization, targeted RL improved the seven-skill mean delta from +1.62 to +3.57 while using 43% less incremental RL compute.