🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Oct 10, 2026
Agents

HarnessSQL: Harness-Native Training for SQL Agents in Realistic Database Environments

First page
HarnessSQL: Harness-Native Training for SQL Agents in Realistic Database Environments
The curator’s take

Haolin Yang, Sirui Han, Yike Guo and colleagues at HKUST (with Microsoft Research, Tsinghua and the University of Macau) propose HarnessSQL, which trains SQL agents inside the same execution harness they use at deployment.

Ask this paper

Key points
01

Problem. Text-to-SQL models are trained to emit static queries, while deployed database agents inspect schemas, run probe queries and revise hypotheses through a harness that is introduced only at inference time.

02

Environments. HarnessSQL builds isolated executable database environments with hidden execution oracles.

03

Training. Teachers are rolled out directly inside the target SQL harness, only verified trajectories are kept for full-sequence SFT, and execution-reward RL follows.

04

Results. On Spider 2.0-SQLite, execution accuracy rises from 15.5% to 45.2% for Qwen3-8B and from 22.2% to 54.8% for Qwen3-14B, with transfer to out-of-distribution interactive benchmarks BIRD-Interact and LiveSQLBench.

Abstract

Text-to-SQL models are commonly trained to map questions directly to static queries, whereas real-world database agents operate through stateful, multi-turn interaction with live databases -- inspecting schemas, executing probe queries, diagnosing errors, and revising hypotheses. This creates a critical train-deploy mismatch, as the execution harness that mediates this interaction is introduced only at inference time. To bridge this gap, we propose HarnessSQL, a harness-native post-training framework that preserves the full interaction structure throughout both supervised fine-tuning and reinforcement learning. HarnessSQL builds isolated, executable database environments paired with hidden execution oracles, rolls out teachers directly inside the target SQL harness, and retains only verified trajectories for full-sequence SFT, followed by execution-reward RL. Across Spider 2.0-SQLite, HarnessSQL dramatically boosts the execution accuracy of compact models, raising Qwen3-8B from 15.5% to 45.2% and Qwen3-14B from 22.2% to 54.8%, while transferring effectively to out-of-distribution interactive benchmarks such as BIRD-Interact and LiveSQLBench. Our findings demonstrate that training database agents directly within their execution harness is essential for mastering complex, long-horizon database workflows.

Every Monday
Get next week’s papers.
Subscribe on Substack