🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Sep 27, 2026
Reasoning · Safety

SAGE: Mitigating Long-Horizon Reasoning Biases via Topological Guidance

First page
SAGE: Mitigating Long-Horizon Reasoning Biases via Topological Guidance
The curator’s take

Xinyue Zeng and colleagues at Virginia Tech, UW-Madison and Dartmouth propose SAGE, which adds structural guidance to long-horizon reasoning to counter exploration and compounding biases under sparse rewards.

Ask this paper

Key points
01

Two biases. Exploration bias pulls models toward locally plausible but unstable branches; compounding bias lets small deviations accumulate with depth and suppress rare rewards.

02

Symbolic Closure Analysis. A theoretical account of how branching structure and sparse reward produce both biases.

03

Two mechanisms. Algebraic sparsification projects candidates onto operator-indexed subspaces to cut spurious branching; hyperbolic embeddings of reasoning states give dense depth-wise signals.

04

Results. Beats baselines across 12 benchmarks and 7 model families, with up to an 8-fold improvement on the Andrews-Curtis problem. Accepted by NeurIPS 2026.

Abstract

Long-horizon reasoning remains a central challenge for large language models (LLMs) under sparse-reward regimes. We argue that this brittleness arises from two biases induced by complex reasoning spaces: an exploration bias, where models are drawn toward locally plausible but structurally unstable branches, and a compounding bias, where small local deviations accumulate across depth and suppress rare rewards. We introduce Symbolic Closure Analysis (SCA) as a theoretical lens characterizing how branching structures and sparse rewards induce these biases in long-horizon reasoning with local admissibility, and as a design principle for structural priors in less formal reasoning tasks. Motivated by this analysis, we propose SAGE (Structural Admissibility-Guided Exploration), a unified framework that injects structural guidance to alleviate exploration bias and compounding bias in long-horizon reasoning. SAGE combines two complementary structural guidance: algebraic sparsification, which projects locally admissible candidates onto operator-indexed algebraic subspaces to suppress spurious branching and mitigate exploration bias, and hyperbolic structural guidance, which embeds reasoning states into a negatively curved space to provide dense depth-wise signals and mitigate compounding bias. Across 12 benchmarks and 7 model families, SAGE outperforms competitive baselines. In particular, SAGE achieves up to an 8-fold improvement on the Andrews-Curtis problem, an open real-world long-horizon task. Code is available at: https://github.com/Susan571/SAGE-NeurIPS2026.

Every Monday
Get next week’s papers.
Subscribe on Substack