🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Oct 8, 2026
Reasoning

SanSi: A Looped Typed Decision Model for System 1.5 Thinking

First page
SanSi: A Looped Typed Decision Model for System 1.5 Thinking
The curator’s take

Shuyu Gan, Young-Jun Lee and Dongyeop Kang at the University of Minnesota present SanSi, a typed decision model that loops the same layers several times before a single option readout, a regime between one-pass decisions and generated reasoning that they call System 1.5 thinking.

Ask this paper

Key points
01

Setup. Typed decision models return a probability for each declared option in one forward pass, with no generated text. SanSi converts a pre-trained looped language model (Ouro-1.4B) into such a model and reads the options after every loop.

02

Training. Every loop is trained with a proper scoring rule, so one model serves any budget from one to eight loops.

03

Results. On 10,027 test decisions from 59 sources, SanSi reaches 72.0% accuracy, 13.5 points above a non-looped model of the same shape trained the same way and 1.8 points below Qwen3.5-4B, which has three times the parameters. Three loops capture 88% of the gain from one to eight loops.

04

Depth generalization. On liar chains of 9 to 16 steps, beyond the 8 seen in training, SanSi is right on 72.6% of items while Qwen3.5-4B is at chance (50.0%).

05

As a judge. Used as a reward judge for RL without gold answers, SanSi raises the generator's F1 by 7.7 points.

Abstract

Typed decision models answer a declared question without generating text: a decision head returns a probability for each of the declared options in a single forward pass. A single pass is fast, intuitive System 1 thinking. We study what lies between one pass and generated reasoning: looping, in which the same layers are recursively applied several times before one typed readout. Each loop lets the model revise its hidden state before it commits to an answer, without generating a token; we call this System 1.5 thinking. We propose SanSi, which turns a pre-trained looped language model into a typed decision model. The option probabilities are read after every loop, and every loop is trained with a proper scoring rule, so that one model serves every budget from one loop to eight in a single run. On 10,027 test decisions from 59 sources, SanSi reaches 72.0% accuracy: 13.5 points above a non-looped model of the same shape trained with the same recipe, 5.3 points above a newer non-looped model of its size, and 1.8 points below one with three times the parameters. On two depth-controlled tasks, loops extend the solvable depth beyond the depths seen in training, where the larger single-pass model fails. Used as the judge for policy optimization with reinforcement learning, without gold answers, SanSi raises the generator's F1 by 7.7 points.

Every Monday
Get next week’s papers.
Subscribe on Substack