🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Sep 29, 2026
Agents

Up and Down the Abstraction Ladder: Code-Based Skills for Language Agents

First page
Up and Down the Abstraction Ladder: Code-Based Skills for Language Agents
The curator’s take

Bartłomiej Cupiał, Jens Tuyls and colleagues from the University of Warsaw, Princeton (Eysenbach, Narasimhan), UCL, Mila and Mistral AI study how giving a language agent a library of code-based skills changes its performance, cost and learning speed, using NetHack as the long-horizon testbed.

Ask this paper

Key points
01

CodeHack library. A set of code-based NetHack skills with natural-language descriptions; the code handles recurring local decisions and the model chooses which skills to call and in what order.

02

Three action interfaces. Agents act with primitives only, skills only, or skills plus primitives, compared under zero-shot prompting, supervised fine-tuning and RL.

03

Zero-shot results. Across 14 models, skills nearly triple average game progression over primitives and cut inference cost per episode by 86%.

04

RL results. Skill-only and mixed controllers gain 7.2x and 8.6x more dungeon depth than primitive-only controllers over the same training budget.

05

Keeping primitives. The mixed interface keeps most of the skill benefit while letting the agent drop to low-level actions when a skill does not cover the situation.

Abstract

Language agents struggle to act and learn in environments that require long sequences of low-level actions. Code-based abstractions can make these agents more productive by letting them invoke reusable skills instead of repeatedly selecting individual actions. The code handles recurring local decisions, while the language model decides which skills to use and how to combine them. Yet abstractions are leaky, and situations beyond a skill's capabilities may require a return to primitive actions. Motivated by this tradeoff between productivity and flexibility, we systematically study how code-based action abstraction affects the performance, inference cost, and learning of language agents. We study this in NetHack, a challenging, long-horizon game environment, using CodeHack, our library of code-based skills with natural-language descriptions. We use this library to compare agents restricted to primitives with those using semantic skills alone or in combination with primitives. We evaluate these agents in three settings: zero-shot prompting, supervised fine-tuning, and reinforcement learning. Across a broad zero-shot evaluation on NetHack, we find that compared with primitives, skills nearly triple game progression, while reducing inference cost per episode by 86%. Combining skills with primitives retains much of this benefit while preserving a path back down to low-level actions. Finally, in RL, we find that skill-based agents learn significantly faster than agents acting on primitives, achieving a 7.2x larger average gain in dungeon level over the same training budget. These results show that a supplied skill library can improve performance, efficiency, and learning, while retaining primitives provides flexibility when the library is insufficient. We release CodeHack together with training and evaluation code.

Every Monday
Get next week’s papers.
Subscribe on Substack