AutoGym: Blueprint-First Generation of Verifiable Agent Gyms

Aarati Andrea Noronha, Kavya Ravikumar and Carly Xiaoyu Lin at Amazon AGI introduce AutoGym, a framework that generates complete RL gyms (task, executable environment and verifier) from a small domain seed or from prior model trajectories.
Ask this paper
Blueprint first. The valid solution space, environment requirements and verification criteria are specified before the environment is built, so every task is solvable by construction instead of being checked afterward by an LLM judge.
Difficulty controls. Explicit parameters set task topology, interaction depth, capability axes, question obfuscation and distractor composition. At low settings AutoGym yields 10% hard tasks; at mid-to-hard settings 39% land in the hard band.
Active curriculum. Performance-informed calibration shifts the parameter distribution as models improve; in the running example, agents that applied a threshold from the wrong policy lead to more tasks with competing policies.
Cost. About $100 to $200 per 50-task batch with 20-way parallelism, with 86% of tasks retained after repair.
Comparison. The Agent World Model baseline is mostly easy for Claude Opus 4.6, while AutoGym produces instances that challenge frontier models in productivity and temporal-reasoning settings.
Abstract
Training agents with reinforcement learning requires a gym, comprising a task, an executable environment in which the task can be attempted, and a verifier that reliably distinguishes success from failure. Constructing such gyms remains manual, expensive, and static. Task sets saturate as models improve and are increasingly exposed to contamination. Synthetic generation offers scale, but single-pass synthesis produces tasks whose difficulty is largely cosmetic. Models comparable in capability solve them despite convoluted phrasing, and correctness must be adjudicated post-hoc by unreliable LLM judges. We present AutoGym, a framework that generates complete gyms (tasks, executable environments, and verifiers) from a minimal domain seed or prior model trajectories. AutoGym introduces three mechanisms. (1) Blueprint-first generation specifies the valid solution space, environment requirements, and verification criteria before the environment is materialized, making solvability a construction prerequisite rather than a property verified after the fact. (2) Explicit generation parameters control task topology, interaction depth, capability axes, question obfuscation, and distractor composition, enabling fine-grained difficulty steering. (3) Active curriculum synthesis uses performance-informed calibration to adjust the distribution over these parameters as model capabilities evolve. Across productivity and temporal-reasoning settings, AutoGym generates gyms spanning the capability spectrum, including instances that challenge frontier models.