Pistis Technical Report

The Pistis team at ByteDance introduces 27B and 9B multimodal models on Qwen3.6 and Qwen3.5, trained with interleaved on-policy distillation and RL, plus an automatic harness optimizer.
Ask this paper
IDRL. Alternates on-policy distillation and RL in one loop instead of running them separately or as a joint loss, for steadier optimization and better credit assignment on long agent trajectories.
Two variants. Pistis-Thinking for multimodal reasoning and Pistis-Agentic for long-horizon planning and tool use, which is strongest on multimodal search.
Pistis-Auto-Harnessing. Improves the agent's inference harness iteratively without changing weights or the interaction budget.
Result. Both scales outperform their base models.
Abstract
We introduce the Pistis model family, comprising 27B- and 9B-parameter multimodal large language models built on Qwen3.6 and Qwen3.5, respectively, and developed through a general and scalable post-training framework. The framework first establishes a strong foundation through large-scale multimodal supervised fine-tuning (SFT). Building on this SFT foundation, we propose Interleaved Distillation and Reinforcement Learning (IDRL), a novel post-training paradigm that tightly integrates on-policy distillation and reinforcement learning within a single training loop. By alternating between the two objectives, rather than optimizing either in isolation or combining them in a static joint loss, IDRL enables more effective knowledge transfer, greater optimization stability, and more precise credit assignment for long-horizon agentic trajectories, leading to stronger performance while mitigating common capability trade-offs. At both model scales, the framework produces two specialized variants: Pistis-Thinking, designed to strengthen deep multimodal reasoning, and Pistis-Agentic, which additionally incorporates agentic trajectory data to support long-horizon planning, iterative reasoning, and tool use. Pistis-Agentic is particularly strong in multimodal search. Both scales outperform their corresponding base models. Beyond model-parameter optimization, we further introduce Pistis-Auto-Harnessing (PAH), a system-level method that automatically improves the agent's inference harness through iterative optimization. Experiments demonstrate that PAH enhances the model performance without updating the model parameters or increasing the interaction budget.