JIT-Agent, a model that writes your agent harness
Free while signed in. Answers cite the passages they came from.

Guibin Zhang, Leo Lu and coauthors train JIT-Agent, a model whose output is an agent harness, synthesizing task-adaptive memory, planning, action protocol and tool orchestration on the fly for any off-the-shelf agentic LLM.
Harness as a machine-generatable artifact: A fixed four-module protocol makes the harness composable and generatable, turning what is normally manual and task-specific into a trainable target.
Three trained behaviors: Customize a harness for the task at hand, repair harnesses for stable execution, and self-evolve by distilling performance signals from an expanding archive of prior configurations.
Model-agnostic and large: DeepSeek-V4-Flash surpasses GPT-5.6 on DeepSearchQA (+9.1) and OdysseyBench (+4.3), and GLM-5.2 gains up to +20.2 points.
Competitive with mature runtimes: Generated harnesses are performance-competitive with OpenCode and Claude Code, and the authors frame harness intelligence as a compounding axis orthogonal to model scaling.
Abstract
Agent capability is not determined by the model alone. The agent harness, encompassing memory management, planning strategy, action protocol, and tool/skill orchestration, can dominate the contribution of the underlying foundation model. Yet harness design remains manual, task-specific, and fundamentally unscalable. We present JIT-Agent, a harness intelligence model trained to synthesize task-adaptive agent harnesses on the fly for arbitrary off-the-shelf agentic LLMs. We formalize the agent harness as a composable, machine-generatable artifact governed by a fixed four-module protocol, and train JIT-Agent to customize harnesses for a given task at hand, repair harnesses for stable and reliable execution, and self-evolve by distilling performance signals from an expanding archive of prior harness configurations. Equipped with JIT-Agent as a harness helper, DeepSeek-V4-Flash surpasses GPT-5.6 on DeepSearchQA (+9.1) and OdysseyBench (+4.3), while the already strong GLM-5.2 gains up to +20.2 points. Across controlled evaluations, JIT-Agent-generated harnesses are performance-competitive with mature agent runtimes such as OpenCode and Claude Code and consistently improve multi-scale model families of DeepSeek V4, Mimo-V2.5, and Qwen3.6. To our knowledge, JIT-Agent is the first model purpose-built for just-in-time harness generation, establishing harness intelligence as a trainable, transferable, and compounding dimension of agent capability orthogonal to model scaling.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack