🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
Agents · Robotics

AgentGym-RL

Free while signed in. Answers cite the passages they came from.

First page
AgentGym-RL
The curator’s take

A modular framework for training LLM agents directly via reinforcement learning across realistic environments, plus a simple schedule, ScalingInter-RL, that lengthens interaction horizons over training to improve stability and performance. Results show a 7B open model can rival or beat larger proprietary systems on web navigation, deep search, games, embodied, and science tasks.

Key points
01

What it is: A unified, decoupled RL stack with three pluggable modules (Environment, Agent, Training) that supports PPO, GRPO, REINFORCE++, and runs across WebArena, Deep Search, TextCraft, BabyAI, and SciWorld.

02

Key idea: ScalingInter-RL starts with short horizons to emphasize exploitation and stable learning, then gradually increases allowed turns to encourage exploration and richer behaviors like planning and reflection.

03

Why it matters: Post-training and test-time compute scale better than model size alone for agentic tasks. A 7B model trained with this framework reaches about 58.6% average success and outperforms much larger baselines.

04

Results snapshot: Web navigation: ScalingInter-7B hits 26.00% overall on WebArena, topping GPT-4o at 16.00. Deep search: 38.25 overall, beating GPT-4o 26.75 and close to strong open baselines; best on NQ at 52.00 and ties TriviaQA at 70.00. Games: 91.00 overall on TextCraft and one of the few with a non-zero at Depth 4 (33.33). Embodied: 96.67 on BabyAI, surpassing o3 and GPT-4o on overall accuracy. Science: 57.00 SOTA on SciWorld, with the 7B RL model also strong at 50.50.

05

Training dynamics: Longer horizons too early can collapse learning; short horizons cap performance. ScalingInter-RL avoids both.

06

Engineering notes: Parallelized browsers, reset hooks, and memory-leak fixes enable reliable long rollouts; a visual UI helps inspect trajectories and failure modes.

07

For practitioners: Prefer GRPO over REINFORCE++ for sparse-reward, long-trajectory agent tasks; curriculum on interaction length offers a simple, robust win; budget compute for post-training and inference sampling before scaling parameters.

08

Web navigation: ScalingInter-7B hits 26.00% overall on WebArena, topping GPT-4o at 16.00.

09

Deep search: 38.25 overall, beating GPT-4o 26.75 and close to strong open baselines; best on NQ at 52.00 and ties TriviaQA at 70.00.

10

Games: 91.00 overall on TextCraft and one of the few with a non-zero at Depth 4 (33.33).

11

Embodied: 96.67 on BabyAI, surpassing o3 and GPT-4o on overall accuracy.

12

Science: 57.00 SOTA on SciWorld, with the 7B RL model also strong at 50.50.

Every Monday
Get next week’s papers.

The same picks and the same summaries, in your inbox. Free, and 176 issues deep.

Subscribe on Substack