Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL

Jinhao Dong, Fuli Luo and colleagues at Xiaomi's LLM Core team, with Renmin, Peking and HKU, introduce GAGAR, which uses an agentic grader to rank test-passing trajectories within each RL rollout group and redistributes advantage toward cleaner, more targeted patches.
Ask this paper
Problem. With binary test rewards, GRPO gives every passing trajectory the same advantage, so a focused fix and a patch with out-of-scope changes are reinforced equally.
Grader. An SFT-trained agent sees all trajectories from a mixed-outcome group in a shared workspace, can inspect code and run checks, and ranks the passing patches on approach, precision, minimality, side effects and codebase conventions.
Sum-preserving redistribution. Lower-ranked passes are downweighted and all passing advantages are rescaled so their total positive advantage is unchanged. Downweighting alone destabilized training.
Results. At step 28, GAGAR reaches 62.2% on DeepSWE v1.1, while the binary-reward baseline had dropped from 58.5% at step 20 to 50.2% and was stopped. GAGAR peaks at 63.4% on DeepSWE and 62.5% on SWE-bench Pro, where the baseline plateaus near 59%.
Efficiency and quality. DeepSWE mean turns fall from 132.3 to 111.6. Audited patch quality rises from 3.70 to 4.03, and the best model scores 62.7% on SWE-bench Pro against 60.5% for GPT-5.6 Sol.
Abstract
Reinforcement learning (RL) for code agents often uses executable tests to provide binary rewards. With these rewards, Group Relative Policy Optimization (GRPO) assigns identical advantages to test-passing trajectories within each rollout group, overlooking differences in implementation quality and adherence to task requirements. This leaves the policy without a learning signal that favors clean, targeted implementations over those containing unnecessary or out-of-scope changes. We introduce GAGAR, a framework for quality-aware credit redistribution in code agent RL. Built on dynamic sampling that retains groups containing both passing and failing trajectories, GAGAR places all trajectories from each group in a shared workspace, where an SFT-trained agentic grader jointly inspects them and ranks the test-passing candidates. Based on this ranking, we downweight lower-ranked trajectories and proportionally rescale the advantages of all test-passing trajectories to restore their original sum. This sum-preserving redistribution retains the relative weights established by quality-based downweighting while shifting credit toward higher-quality implementations. We evaluate GAGAR at industrial scale using pre-RL SFT checkpoints of MiMo-V2.6-Flash (310B total parameters) and MiMo-V2.6-Pro (1.02T total parameters). Controlled code-only Flash experiments show improved code agent performance, reduced trajectory-length growth, and more stable training. We further apply GAGAR in large-scale mixed-task RL with both Flash and Pro. Our results support combining test-based verification with groupwise agentic grading to improve the quality and stability of code agent RL.