🚀NEW LABGetting Started with Claude AgentsStart lab
Agents · Robotics

AgentGym-RL

First page
AgentGym-RL
Paper summary

A modular framework for training LLM agents directly via reinforcement learning across realistic environments, plus a simple schedule, ScalingInter-RL, that lengthens interaction horizons over training to improve stability and performance. Results show a 7B open model can rival or beat larger proprietary systems on web navigation, deep search, games, embodied, and science tasks.

Ask this paper

Key points
01

What it is: A unified, decoupled RL stack with three pluggable modules (Environment, Agent, Training) that supports PPO, GRPO, REINFORCE++, and runs across WebArena, Deep Search, TextCraft, BabyAI, and SciWorld.

02

Key idea: ScalingInter-RL starts with short horizons to emphasize exploitation and stable learning, then gradually increases allowed turns to encourage exploration and richer behaviors like planning and reflection.

03

Why it matters: Post-training and test-time compute scale better than model size alone for agentic tasks. A 7B model trained with this framework reaches about 58.6% average success and outperforms much larger baselines.

04

Results snapshot: Web navigation: ScalingInter-7B hits 26.00% overall on WebArena, topping GPT-4o at 16.00. Deep search: 38.25 overall, beating GPT-4o 26.75 and close to strong open baselines; best on NQ at 52.00 and ties TriviaQA at 70.00. Games: 91.00 overall on TextCraft and one of the few with a non-zero at Depth 4 (33.33). Embodied: 96.67 on BabyAI, surpassing o3 and GPT-4o on overall accuracy. Science: 57.00 SOTA on SciWorld, with the 7B RL model also strong at 50.50.

05

Training dynamics: Longer horizons too early can collapse learning; short horizons cap performance. ScalingInter-RL avoids both.

06

Engineering notes: Parallelized browsers, reset hooks, and memory-leak fixes enable reliable long rollouts; a visual UI helps inspect trajectories and failure modes.

07

For practitioners: Prefer GRPO over REINFORCE++ for sparse-reward, long-trajectory agent tasks; curriculum on interaction length offers a simple, robust win; budget compute for post-training and inference sampling before scaling parameters.

08

Web navigation: ScalingInter-7B hits 26.00% overall on WebArena, topping GPT-4o at 16.00.

09

Deep search: 38.25 overall, beating GPT-4o 26.75 and close to strong open baselines; best on NQ at 52.00 and ties TriviaQA at 70.00.

10

Games: 91.00 overall on TextCraft and one of the few with a non-zero at Depth 4 (33.33).

11

Embodied: 96.67 on BabyAI, surpassing o3 and GPT-4o on overall accuracy.

12

Science: 57.00 SOTA on SciWorld, with the 7B RL model also strong at 50.50.

Every Monday
Get next week’s papers.
Subscribe on Substack