🚀NEW LABGetting Started with Claude AgentsStart lab
Agents · Evaluation

Harnesses Are Not Uniformly Better

First page
Harnesses Are Not Uniformly Better
Paper summary

This paper studies LLM agent harnesses through the lens of inference-time trajectory alignment, separating a harness into two mechanisms: task decomposition, which structures a task into sub-goals, and guided execution, which reshapes local action distributions during execution. The key finding is that more elaborate harnesses are not uniformly better. Increasing decomposition or guidance can improve execution but can also reduce final task success, producing concrete failure modes like over-decomposition, over-pruning, and hallucinated execution. Strikingly, partial harnesses that specify only the initial steps and leave the rest to the agent can reach a higher pass rate than fully structured workflows.

Ask this paper

Every Monday
Get next week’s papers.
Subscribe on Substack