Shreya Rajpal and Ammon explain how Nubank uses Snow Globe’s simulation platform to generate synthetic multi-turn conversations, enabling 20× faster evaluation and deployment of AI customer support agents by reducing reliance on slow, risky production data. This simulation-driven approach not only accelerates iteration and innovation but also improves agent reliability and customer satisfaction by closely aligning simulated evaluations with real-world performance.
In this talk, Shreya Rajpal, CEO of Snow Globe, and Ammon, Principal Machine Learning Engineer at Nubank, discuss how Nubank leverages simulations to accelerate the deployment of AI agents for customer support by 20 times. Nubank, a leading digital bank in Latin America with over 135 million customers, uses AI agents alongside human experts to handle routine customer inquiries efficiently while humans focus on complex cases. The key insight shared is that generating evaluation data through simulations rather than waiting for production data significantly speeds up the development and deployment cycle of AI agents.
The speakers emphasize the critical role of evaluations (evals) in building effective AI agents, highlighting that while metrics can be systematically developed and aligned with human judgment, obtaining high-quality evaluation data remains a major bottleneck. Traditional methods of collecting eval data—manual authoring and production traces—are either time-consuming or risky, as testing in production can negatively impact real users. Moreover, the complexity of multi-turn agent interactions makes data generation and annotation expensive and slow, further hindering rapid experimentation and iteration.
Simulations provide a powerful solution by creating synthetic, multi-turn conversations that mimic real user interactions with agents. Snow Globe’s platform enables Nubank to wrap their agents and simulate diverse user personas and scenarios without requiring code changes. These simulations generate large datasets with consistent state and tool mocks, allowing for rapid offline evaluation of agents. This approach drastically reduces the time needed for offline evaluations from days or weeks to mere hours or minutes, enabling faster iteration and more frequent A/B testing without risking customer experience.
The talk also highlights the strong correlation between simulation-generated data and real production data, validated through human expert reviews and evaluation metrics. Simulations have helped Nubank catch regressions and inefficiencies before deployment, improving both customer satisfaction (TNPS) and self-service rates. The ability to quickly test different open-source models within the agent harness using simulations has saved weeks of effort and accelerated innovation. This has made Nubank’s AI agent development more efficient, reliable, and scalable.
In conclusion, the presenters outline three key takeaways: first, generating evaluation data via simulations can dramatically shorten agent release cycles; second, it is essential to close the gap between simulation and real-world data through rigorous validation to build trust in simulation results; and third, the future of self-improving AI agents depends on having aligned metrics and reliable data generation methods. By adopting simulation-based evaluation, enterprises can unlock continuous improvement loops for AI agents, driving better customer experiences and operational efficiency.