🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Sep 24, 2026
Reinforcement Learning · Code · Agents

FLARE: A Full-Lifecycle Dense Supervision Paradigm for Long-Horizon Coding Agents via Generative Reward Model

First page
FLARE: A Full-Lifecycle Dense Supervision Paradigm for Long-Horizon Coding Agents via Generative Reward Model
The curator’s take

Jingxuan Xu, Gang Wu, Yanan Wu, Yutao Mou and colleagues (independent researchers with Peking, Nanjing and BUPT) propose FLARE, which trains a lightweight generative reward model to give step-level risk feedback to long-horizon coding agents during inference, SFT and RL.

Ask this paper

Key points
01

RADAR diagnosis. An offline causal-chain backtracking procedure extracts step-level supervision from trajectories without hindsight leakage and distills it into the generative reward model.

02

Inference as active scaffold. The reward model flags high-risk steps and FLARE re-executes from that breakpoint instead of rolling out whole new trajectories.

03

Compute result. FLARE with one rollout outperforms global rollout with five while using 5x fewer tokens.

04

Training signals. The same scores rerank SFT data and serve as dense step rewards in RL, addressing sparse pass/fail rewards.

05

Training gains. Process-aware data curation gives a 19.13% relative gain in SFT, and dense rewards give a 9.19% relative gain in RL.

Abstract

While test-time scaling enhances Large Language Model (LLM) agents in long-horizon software engineering (SWE), sparse binary rewards (Pass/Fail) create a severe credit assignment crisis and waste failed exploratory trajectories. Current trajectory optimization and scaling methods are costly and structurally limited, relying on heuristic state reuse without causal diagnosis or delayed scalar scoring without actionable online guidance. We propose FLARE (Full-Lifecycle Alignment and Reward Engine), a novel dense supervision paradigm driven by a lightweight Generative Reward Model (GRM). First, RADAR, an offline causal-aware diagnostic framework, extracts high-fidelity, hindsight-free supervision through causal-chain backtracking to distill a GRM providing real-time, step-level risk feedback. Second, FLARE uses this GRM to continuously optimize the agent across its entire lifecycle. During inference, FLARE acts as an Active Scaffold, autonomously intercepting high-risk generation steps for localized breakpoint re-execution, drastically reducing compute overhead. During post-training, the GRM's structured signals serve as process-supervised reranking scores for Supervised Fine-Tuning (SFT) and step-level dense rewards for Reinforcement Learning (RL), mitigating policy collapse in sparse environments. Extensive evaluations show that FLARE establishes a new Pareto frontier across the agent lifecycle: FLARE (N=1) outperforms Global Rollout (N=5) with a 5x reduction in token consumption. Extending FLARE to training overcomes the sparse reward problem in long-horizon interactive tasks, delivering relative performance gains of 19.13% in SFT through process-aware data curation and a consistent 9.19% improvement in RL.

Every Monday
Get next week’s papers.
Subscribe on Substack