Kimi K2.5: Visual Agentic Intelligence
Free while signed in. Answers cite the passages they came from.

Kimi K2.5 is an open-source multimodal agentic model from Moonshot AI that jointly optimizes text and vision capabilities through native multimodal pretraining on 15 trillion mixed tokens, zero-vision SFT, and joint reinforcement learning. K2.5 also introduces Agent Swarm, a parallel agent orchestration framework that dynamically decomposes complex tasks into concurrent subtasks, reducing latency by up to 4.5x over single-agent baselines. - **Joint text-vision optimization:** K2.5 uses early fusion with a lower vision ratio during pretraining (rather than late-stage heavy vision injection), achieving better results across both modalities. A key finding is that zero-vision SFT - using only text SFT data - is sufficient to activate visual reasoning and tool use, while visual RL actually improves text benchmarks like MMLU-Pro (+1.7%) and GPQA-Diamond (+2.1%). - **Agent Swarm with Parallel-Agent RL:** The framework trains a learnable orchestrator via RL to decompose tasks and delegate subtasks to frozen specialized subagents running in parallel. This decoupled design avoids credit assignment ambiguity, and improves item-level F1 from 72.8% to 79.0% on wide-search scenarios while significantly reducing inference latency. - **State-of-the-art agentic performance:** K2.5 achieves 74.9% on BrowseComp (with context management), 77.1% on DeepSearchQA, and 57.4% on Seal-0, outperforming GPT-5.2, Claude Opus 4.5, and Gemini 3 Pro. It also scores 96.1% on AIME 2025, 76.8% on SWE-Bench Verified, and establishes new records in long-video comprehension. - **Token-efficient RL with Toggle:** K2.5 introduces Toggle, a training heuristic that alternates between budget-constrained and standard scaling phases during RL, reducing output tokens by 25-30% with negligible performance impact while maintaining strong test-time scaling capabilities.
Joint text-vision optimization: K2.5 uses early fusion with a lower vision ratio during pretraining (rather than late-stage heavy vision injection), achieving better results across both modalities. A key finding is that zero-vision SFT - using only text SFT data - is sufficient to activate visual reasoning and tool use, while visual RL actually improves text benchmarks like MMLU-Pro (+1.7%) and GPQA-Diamond (+2.1%).
Agent Swarm with Parallel-Agent RL: The framework trains a learnable orchestrator via RL to decompose tasks and delegate subtasks to frozen specialized subagents running in parallel. This decoupled design avoids credit assignment ambiguity and improves item-level F1 from 72.8% to 79.0% on wide-search scenarios while significantly reducing inference latency.
State-of-the-art agentic performance: K2.5 achieves 74.9% on BrowseComp (with context management), 77.1% on DeepSearchQA, and 57.4% on Seal-0, outperforming GPT-5.2, Claude Opus 4.5, and Gemini 3 Pro. It also scores 96.1% on AIME 2025, 76.8% on SWE-Bench Verified, and establishes new records in long-video comprehension.
Token-efficient RL with Toggle: K2.5 introduces Toggle, a training heuristic that alternates between budget-constrained and standard scaling phases during RL, reducing output tokens by 25-30% with negligible performance impact while maintaining strong test-time scaling capabilities.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack