Bootstrapping Conversational Recommendation Agents At Spotify: Synthetic Data Generation and Self-Improvement Loops

Enrico Palumbo and colleagues at Spotify (RecSys 2026) describe how they built a multi-turn conversational recommendation agent before launch, when no real conversation data existed, using synthetic conversations and an automated prompt-improvement loop.
Ask this paper
Synthetic multi-turn data. A pipeline expands single-turn prompts into realistic multi-turn user-agent conversations; human annotation scored them above 90% on every quality dimension.
Self-improvement loop. Variance-based contrastive optimization finds unstable planning and tool-use behavior, and a coding agent revises the agent prompt; humans review every change before deployment.
Offline gain. The loop adds +8% quality on top of a heavily hand-tuned production prompt.
Online A/B test. Against the previous refinement-only experience: +14% listening, +5% weekly active users and 5% fewer skips.
Where agents still fail. Instruction retention and content refinement over evolving lists remain hard for frontier models, and quality drops once conversations pass four turns.
Abstract
Conversational recommendation agents are a new paradigm for content discovery, enabling users to express complex intents through natural language (e.g., "recommend Italian indie artists I haven't heard before"). A central challenge in building such agents is optimizing agent planning -- deciding how to select, sequence, and invoke tools -- particularly in cold-start settings where real user interactions are not yet available. We introduce a pipeline for multi-turn synthetic data generation and a self-improvement loop to address this challenge. The synthetic data pipeline transforms single-turn prompts into realistic multi-turn conversations, enabling systematic evaluation before launch. The self-improvement loop combines variance-based contrastive optimization with iterative refinement through a coding agent, automatically identifying and fixing planning and tool-use errors. Our approach improves quality by +8% on top of a highly optimized manual prompt. The system has been productionized and significantly accelerated iteration cycles for the launch of a conversational recommendation agent at Spotify. Online A/B tests demonstrate its effectiveness, with +14% user listening, +5% increase in weekly active users, and a 5% reduction in skip rate compared to a prior experience supporting only session refinement. This work provides a practical framework for accelerating the development of conversational recommendation agents in industry.