Bogdan Gaza, CTO of Duality, shares insights on overcoming engineering challenges to scale synthetic data generation to trillions of tokens using their BeyondWeb paraphrasing method, which enables smaller models to achieve high accuracy efficiently. He emphasizes the importance of unified infrastructure, optimized resource management, and inference tuning to streamline large-scale synthetic data production across diverse domains.
In his talk, Bogdan Gaza, co-founder and CTO of Duality, discusses the engineering challenges and lessons learned from scaling synthetic data generation to the scale of trillions of tokens. He begins by explaining the necessity of synthetic data due to the “data wall” problem, where acquiring exponentially more real-world data becomes impractical for improving model performance. Synthetic data, especially high-quality paraphrased data generated from existing datasets, helps overcome the limitations of internet-sourced data and enables training large models more efficiently.
Bogdan introduces Duality’s synthetic data recipe called BeyondWeb, which focuses on paraphrasing high-quality data points to create vast datasets. He presents evidence that models trained with BeyondWeb synthetic data can achieve comparable or better accuracy than those trained on larger real datasets, often requiring fewer tokens and computational resources. This approach allows smaller models to perform on par with larger ones, demonstrating the efficiency and effectiveness of their synthetic data generation method.
The talk then shifts to the technical infrastructure behind generating synthetic data at scale. Duality uses Kubernetes clusters orchestrated with tools like Ray, Spark, and vLLM to manage workloads across CPU and GPU resources, including specialized H100 GPU clusters. They emphasize the importance of unifying research and engineering workflows on a single infrastructure to streamline experimentation, training, and evaluation processes, enabling faster and more reliable synthetic data generation.
Bogdan highlights several key engineering challenges encountered at scale, such as metadata management bottlenecks when accessing millions of data partitions from S3 storage, GPU instability causing task failures, and the complexity of orchestrating tasks across multiple clusters with different resource requirements. Solutions include batching metadata requests to reduce processing time from days to hours, implementing checkpointing and smaller data partitions to minimize lost work from failures, and carefully scheduling CPU and GPU resources to optimize cluster utilization.
Finally, Bogdan discusses the importance of inference optimization, particularly tuning parameters in vLLM to improve throughput by up to 40%. He summarizes the overall lessons learned: the need for robust failover systems, atomic resource scheduling, and continuous parameter tuning to handle synthetic data generation at the trillion-token scale. Duality has successfully scaled from generating billions to trillions of synthetic tokens, supporting diverse domains like web, math, code, and legal data, and continues to refine their methods while actively hiring experts to expand their capabilities.
Useful Links
- BeyondWeb synthetic data recipe paper — Directly explains the synthetic data recipe central to the talk’s claims about synthetic data quality and scaling.
- DatologyAI BeyondWeb blog post — Provides detailed explanation and results of the synthetic data recipe central to the talk.
- Ray distributed execution framework — Key software infrastructure enabling the synthetic data generation pipeline at scale.
- vLLM large language model serving system — Central to inference tuning and throughput improvements discussed in the video.
- Kubernetes container orchestration system — Fundamental infrastructure component for managing compute resources in the synthetic data pipeline.