🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Oct 5, 2026
Reasoning

RLTL;DR: Self-improvement by Internalizing Self-generated Feedback

First page
RLTL;DR: Self-improvement by Internalizing Self-generated Feedback
The curator’s take

Michael Kirchhof, Andrew Szot, Alexander Toshev and colleagues at Apple introduce RLTL;DR, an RL recipe for tasks where the policy never succeeds: after each failed attempt the policy writes a one-line insight from the verifier output, later rollouts are conditioned on those insights, and the insight tokens are trained on so the lesson moves into the weights.

Ask this paper

Key points
01

Setting. Tool-calling and coding tasks filtered to Pass@128 = 0 for a Qwen 3.5 9B Thinking policy, so standard GRPO receives no positive reward and stays flat at 0% to 1% Pass@1.

02

Sequential rollouts. Each failed attempt yields a TL;DR insight of about 17 tokens (for example "Remember to paginate search results"), and the next attempt sees all previous insights. This alone finds a solution for 14% to 59% of tasks during exploration, but test-time Pass@1 without insights barely moves.

03

Training on the insights. Turning on the backpropagation mask for the in-context insight tokens is the change that matters. RLTL;DR reaches 14% to 31% Pass@1 with insights in context and 12% to 13% with no insight at eval time.

04

SFTL;DR. Training only on 4k (task, insight) pairs, with no rollouts shown or trained on, recovers almost all of RLTL;DR's gain and matches classical SFT on full rollouts.

05

Why it matters. It gives a way to learn on tasks beyond a model's current reach without a teacher model or reference solutions.

Abstract

The common paradigm of reinforcement learning with verifiable rewards (RLVR) is to let agents make multiple attempts at a task, and optimize towards the successful ones. This becomes problematic in the realms of self-improvement, where tasks are so difficult that the agent has a low or even no chance of success, and where there are no teacher models or example solutions to distill from. In this paper, we introduce RLTL;DR. After each failed attempt, we show the policy the verifier outputs and let it write its own feedback, in the form of a single TL;DR insight. The next rollout is conditioned on all previous insights, and we sequentially sample rollouts until a solution is found. Moreover, we enable backpropagation on the in-context insights to internalize a direct task to insight mapping. On challenging tool-calling and coding datasets (filtered to Pass@128=0), standard GRPO training of a Qwen 3.5 9B Thinking policy stays flat at a Pass@1 of 0% to 1%. RLTL;DR breaks through this learning barrier, achieving a Pass@1 of 14-31% with insights in context during training and, crucially, 12-13% when no insight is in context at eval time. We identify that the key is the task to insight internalization. To study this further, we reduce our approach to SFTL;DR, training only on (task, insight) tuples, without showing or backpropagating on any rollouts. Training on only 4k of these tuples recovers almost the full performance of RLTL;DR and classical SFT on full rollouts. This demonstrates a promising compacted training paradigm of the form "on this sort of task, keep this sort of thing in mind", which we hope to inspire future research on.

Every Monday
Get next week’s papers.
Subscribe on Substack