HarnessSQL: Harness-Native Training for SQL Agents in Realistic Database Environments

Haolin Yang, Sirui Han, Yike Guo and colleagues at HKUST (with Microsoft Research, Tsinghua and the University of Macau) propose HarnessSQL, which trains SQL agents inside the same execution harness they use at deployment.
Ask this paper
Problem. Text-to-SQL models are trained to emit static queries, while deployed database agents inspect schemas, run probe queries and revise hypotheses through a harness that is introduced only at inference time.
Environments. HarnessSQL builds isolated executable database environments with hidden execution oracles.
Training. Teachers are rolled out directly inside the target SQL harness, only verified trajectories are kept for full-sequence SFT, and execution-reward RL follows.
Results. On Spider 2.0-SQLite, execution accuracy rises from 15.5% to 45.2% for Qwen3-8B and from 22.2% to 54.8% for Qwen3-14B, with transfer to out-of-distribution interactive benchmarks BIRD-Interact and LiveSQLBench.
Abstract
Text-to-SQL models are commonly trained to map questions directly to static queries, whereas real-world database agents operate through stateful, multi-turn interaction with live databases -- inspecting schemas, executing probe queries, diagnosing errors, and revising hypotheses. This creates a critical train-deploy mismatch, as the execution harness that mediates this interaction is introduced only at inference time. To bridge this gap, we propose HarnessSQL, a harness-native post-training framework that preserves the full interaction structure throughout both supervised fine-tuning and reinforcement learning. HarnessSQL builds isolated, executable database environments paired with hidden execution oracles, rolls out teachers directly inside the target SQL harness, and retains only verified trajectories for full-sequence SFT, followed by execution-reward RL. Across Spider 2.0-SQLite, HarnessSQL dramatically boosts the execution accuracy of compact models, raising Qwen3-8B from 15.5% to 45.2% and Qwen3-14B from 22.2% to 54.8%, while transferring effectively to out-of-distribution interactive benchmarks such as BIRD-Interact and LiveSQLBench. Our findings demonstrate that training database agents directly within their execution harness is essential for mastering complex, long-horizon database workflows.