ComputerRL
Free while signed in. Answers cite the passages they came from.

A framework for autonomous desktop agents that unifies API calls with GUI actions, plus a scalable RL stack and a training recipe (Entropulse) that alternates RL and SFT to sustain exploration. Evaluated on OSWorld, it sets a new SOTA with strong gains in efficiency and robustness.
APIāGUI action space. Moves beyond humanācentric GUIs by combining programmatic APIs with direct GUI control. LLMs help autoāgenerate appāspecific APIs via requirement analysis, implementation, and unit tests, lowering the cost of adding new tools.
Massively parallel desktop env. A refactored Ubuntu VM cluster (qemuāinādocker + gRPC) delivers thousands of concurrent instances with improved stability, monitoring, and AgentBenchācompatible interfaces, enabling largeāscale online RL.
Fully asynchronous RL. Built on AgentRL with decoupled actors/trainers, dynamic batching, and bounded replay to reduce offāpolicy bias and maximize GPU utilization during longāhorizon desktop rollouts.
Entropulse training. After an initial stepālevel GRPO phase with verifiable, ruleābased rewards, successful rollouts are distilled via SFT to restore entropy, then RL resumes, yielding higher rewards and sustained improvements.
Results and analysis. AUTOGLMāOSā9B reaches 48.1% on OSWorld and 47.3% on OSWorldāVerified, outperforming OpenAI CUA o3, UIāTARSā1.5, and Claude Sonnet 4, while using up to oneāthird the action steps of strong baselines; ablations show APIāGUI and multiāstage training drive the gains. Error sources cluster into multiāapp coordination, vision, and operational slips.
Get next weekās papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack