Visualization-of-Thought
Free while signed in. Answers cite the passages they came from.

Microsoft's Visualization-of-Thought (VoT) prompts LLMs to emit intermediate "mental images" of their reasoning state, lifting spatial-reasoning accuracy on grid-world tasks and beating multimodal baselines that actually see images.
Mental-imagery prompting: The model renders each reasoning step as an ASCII/grid-style visualization, which is then fed back in to constrain subsequent steps, echoing human mental imagery.
Benchmarks: VoT is evaluated on natural-language navigation, visual navigation, and visual tiling in 2D grid worlds, targeting multi-hop spatial reasoning.
Beats multimodal LLMs: Text-only LLMs with VoT outperform contemporary multimodal LLMs on the same tasks, showing explicit state visualization can substitute for actual image tokens.
NeurIPS 2024: The method was accepted to NeurIPS 2024 and positions mental imagery as a general-purpose tool for strengthening reasoning in otherwise text-only models.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack