🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Sep 25, 2026
Agents

Making Agents More Consistent: Skills Should Form Habits for Repeat Tasks

First page
Making Agents More Consistent: Skills Should Form Habits for Repeat Tasks
The curator’s take

Travis Weber and Rohit Taneja (Pheo) measure how inconsistent agents are on repeated work and propose skill habit formation, where an agent turns recurring plans from its own history into deterministic scripts that are admitted only after passing a series of gates.

Ask this paper

Key points
01

The inconsistency. Running 42 tasks three times each, 38% to 74% of tasks returned disagreeing answers depending on the model, and 95.3% to 97.2% of generated tokens went to re-deriving a plan the system already had.

02

Habit candidates. The agent mines its execution history for deterministic variants that compete with the incumbent; each declares the input region it handles, and inputs outside it fall through to normal reasoning.

03

Four gates. Candidates pass four gates of increasing cost, the central one comparing the execution trace to a retained reference within that reference's own run-to-run tolerance.

04

Text-to-SQL results. Reasoning arms reproduced their own output on 11 to 26 of 42 repeated questions, while the habit-formed variant reproduced on all 456 repeated dispatches, was non-inferior to every arm it replaced and used 14% to 56% fewer tokens.

05

Measured cost. The guard wrongly admitted 2.6% of natural paraphrases and 26% of near-boundary inputs, and separating routing from parameter extraction raised accuracy from 0.888 to 0.952 at 43% of the cost.

Abstract

On repeated work, agents are inconsistent. We ran 42 tasks three times each and found that, depending on the model, 38% to 74% returned answers that did not agree. Consistency is what a buyer, an auditor, or a regulator requires, and agents do not have it. They are wasteful too: 95.3% to 97.2% of what an agent generates goes to re-deriving a plan the system already knows. We propose skill habit formation. An agent mines its own execution history for candidate skills, deterministic variants that compete against the incumbent rather than replacing it. A candidate declares the region of input space it claims, so the common case runs as a script and the rest falls through to reasoning. Four gates of ascending cost admit candidates; the central one tests a candidate's execution trace against a retained reference, within a tolerance measured from that reference's own run-to-run variability. On text-to-SQL, three of four reasoning arms reproduced their own output on 11 to 13 of 42 repeated questions and the fourth on 26 of 42, while a habit-formed variant reproduced on all 456 dispatches we repeated and was non-inferior to every arm it replaced (p<0.0001). It also used 14% to 56% fewer tokens, turning net positive after 7 to 53 reuses. We measured what this costs in accuracy. The guard admitted work it should have deferred on 2.6% of natural paraphrases and 26% of inputs near its boundary, and 11 of 13 such failures were invisible to the trace-conformance gate at any threshold. Deterministic errors repeat exactly: a bad habit is as reliable as a good one, and that is the price of the property that makes the system auditable. Separating routing from parameter extraction raised end-to-end accuracy from 0.888 to 0.952 at 43% of the cost.

Every Monday
Get next week’s papers.
Subscribe on Substack