🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Sep 27, 2026
Agents

Screen Before You Serve: Simulation for Production Customer Experience AI Agents at 140M Scale

First page
Screen Before You Serve: Simulation for Production Customer Experience AI Agents at 140M Scale
The curator’s take

A Nubank team with Snowglobe describes simulation-based screening of customer-experience agents before deployment on Nubank's highest-volume chat-support agent in Brazil.

Ask this paper

Key points
01

Workflow. Synthetic customers react to agent replies and simulated tool outputs support multi-step flows without touching production backends.

02

Correlation. Across four deployed versions, simulated and production version-level evaluator scores correlate highly.

03

Live gains. Simulation-guided iteration raised transactional NPS by 36.69 points in a live A/B test.

04

Model screening. After screening open-weight configurations in over 16,000 simulated conversations, the selected model raised self-service rate by 8.82 points with no significant tNPS change.

Abstract

Customer experience (CX) agents use tools and large language models to address customer requests and guide conversational interactions with an organization's products. Improving these agents, especially in regulated industries, is difficult: they must detect intent, follow complex operational policies and use tools reliably. Manual end-to-end testing offers limited coverage, while live experiments expose customers to failures that can erode trust. We present a hypothesis-driven simulation workflow for screening candidate CX agents before deployment. Synthetic customers react to agent responses and simulated tool outputs enable multi-step agentic workflows without invoking production backends. We use the Snowglobe simulator on Nubank's Card Delivery agent and its expanded successor, Card Management - Nubank's highest-volume chat-support agent in Brazil. Across 4 deployed versions, simulated and production version-level binary evaluator scores show high correlation. Simulation-guided iteration increased transactional net promoter score (tNPS) by 36.69 points in a live A/B test. We also screened open-weight configurations in over 16,000 simulated conversations. In a subsequent live A/B test, the selected model increased self-service rate (SSR) by 8.82 percentage points to the highest level observed at Nubank, with no statistically significant change in tNPS. Simulation made broad exploration of models, reasoning settings, and prompts feasible without customer exposure, enabling production improvements that would have been impractical to pursue through live experimentation alone.

Every Monday
Get next week’s papers.
Subscribe on Substack