SanSi: A Looped Typed Decision Model for System 1.5 Thinking

Shuyu Gan, Young-Jun Lee and Dongyeop Kang at the University of Minnesota present SanSi, a typed decision model that loops the same layers several times before a single option readout, a regime between one-pass decisions and generated reasoning that they call System 1.5 thinking.
Ask this paper
Setup. Typed decision models return a probability for each declared option in one forward pass, with no generated text. SanSi converts a pre-trained looped language model (Ouro-1.4B) into such a model and reads the options after every loop.
Training. Every loop is trained with a proper scoring rule, so one model serves any budget from one to eight loops.
Results. On 10,027 test decisions from 59 sources, SanSi reaches 72.0% accuracy, 13.5 points above a non-looped model of the same shape trained the same way and 1.8 points below Qwen3.5-4B, which has three times the parameters. Three loops capture 88% of the gain from one to eight loops.
Depth generalization. On liar chains of 9 to 16 steps, beyond the 8 seen in training, SanSi is right on 72.6% of items while Qwen3.5-4B is at chance (50.0%).
As a judge. Used as a reward judge for RL without gold answers, SanSi raises the generator's F1 by 7.7 points.
Abstract
Typed decision models answer a declared question without generating text: a decision head returns a probability for each of the declared options in a single forward pass. A single pass is fast, intuitive System 1 thinking. We study what lies between one pass and generated reasoning: looping, in which the same layers are recursively applied several times before one typed readout. Each loop lets the model revise its hidden state before it commits to an answer, without generating a token; we call this System 1.5 thinking. We propose SanSi, which turns a pre-trained looped language model into a typed decision model. The option probabilities are read after every loop, and every loop is trained with a proper scoring rule, so that one model serves every budget from one loop to eight in a single run. On 10,027 test decisions from 59 sources, SanSi reaches 72.0% accuracy: 13.5 points above a non-looped model of the same shape trained with the same recipe, 5.3 points above a newer non-looped model of its size, and 1.8 points below one with three times the parameters. On two depth-controlled tasks, loops extend the solvable depth beyond the depths seen in training, where the larger single-pass model fails. Used as the judge for policy optimization with reinforcement learning, without gold answers, SanSi raises the generator's F1 by 7.7 points.