🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
← All papers  /  Sep 20, 2026
Robotics

MAGMA-GEN: Validated Recovery Supervision from Ambiguous Failures via Counterfactual Re-Execution

First page
MAGMA-GEN: Validated Recovery Supervision from Ambiguous Failures via Counterfactual Re-Execution
The curator’s take

Loan Bernat, Matthieu Grard, Ariane Herbulot and Florent Lamiraux turn ambiguous failed robot rollouts into recovery supervision by re-executing proposed corrections and keeping only the ones that actually help.

Ask this paper

Key points
01

Why failed rollouts are ambiguous. A poor downstream state may come from an invalid high-level decision, partial observation, or a valid decision whose stochastic low-level skill failed, and nothing in the trajectory distinguishes them.

02

Privileged coach, then verification. A privileged coach hypothesizes an early decision-level error and proposes a localized correction, and because that diagnosis is fallible, the candidate is kept only if re-execution from the same state under matched conditions improves downstream progress.

03

Supervision from the agent's own failures. The pipeline produces training examples drawn from the policy's actual failure distribution without per-step human demonstrations.

04

Results. On interactive long-horizon manipulation, task success and recovery both improve over distillation and trajectory-repair baselines, in simulation and on a real robot, under evolving task constraints.

05

Significance. The counterfactual re-execution step is what converts a guess about the cause of failure into verified data, which is transferable to any agent setting with a resettable environment.

Abstract

Hierarchical robotic systems executing long-horizon manipulation tasks must make high-level semantic decisions that orchestrate stochastic low-level skills. In this setting, failed rollouts are ambiguous: a poor downstream state may reflect an invalid high-level decision, partial observation, or a valid decision whose physical execution failed. Traditional supervised learning lacks data for such recovery states, while reinforcement learning struggles with sparse rewards and non-local credit assignment. We propose MAGMA-GEN, an on-policy data-generation pipeline that converts ambiguous failed rollouts into validated recovery supervision. MAGMA-GEN first uses a privileged coach to hypothesize an early decision-level error and propose localized correction or recovery actions. Because this diagnosis is fallible, candidates are retained only if re-execution from the same state under matched conditions improves downstream progress. This produces supervised examples from the agent's own failure distribution without per-step human demonstrations. Evaluated on interactive long-horizon manipulation tasks, MAGMA-GEN improves task success and recovery capabilities, against distillation and trajectory-repair baselines under evolving task constraints in both simulation and real-robot execution.

Every Monday
Get next week’s papers.
Subscribe on Substack