Silicon Valley data labeling startups are supplying crucial AI training data to Chinese tech giants, enabling China’s AI industry to advance rapidly despite US restrictions on advanced AI chip exports. This unregulated flow of high-quality training data, embedding proprietary expertise, highlights a significant and overlooked facet of the US-China AI competition and underscores the need for broader export controls beyond hardware.
Silicon Valley’s data labeling startups, which supply training data to leading AI companies like OpenAI and Anthropic, are increasingly serving China’s rapidly growing AI industry. At the International Conference on Machine Learning in Seoul, US sales teams discovered strong demand from Chinese AI firms, including companies like Tencent and 10 Cent, for specialized training data across sectors such as finance, cybersecurity, and advanced AI research areas. Despite US restrictions on selling advanced AI chips to China, there are no similar controls on the sale of training data, which is crucial for developing sophisticated AI models.
Training data, which involves curated human expertise shaped by detailed task designs and quality controls, is a vital component in teaching AI models to perform complex professional tasks like financial modeling and coding. This data infrastructure is a lucrative business generating hundreds of millions of dollars annually and is helping Chinese AI labs narrow the performance gap with their American counterparts. Experts emphasize that after computational power, high-quality data is the most critical factor in advancing AI capabilities, especially in challenging domains.
Chinese AI labs employ several strategies to accelerate their progress, including recruiting researchers from abroad, using outputs from AI models like ChatGPT to train their own systems, and purchasing the same high-quality training data sets from US vendors. The latter approach allows Chinese companies to acquire not just raw data but also the proprietary judgment embedded in the data specifications and rubrics developed by Silicon Valley professionals. This effectively transfers key elements of AI expertise from the US to China through commercial data transactions.
Communications reviewed by Forbes reveal that many top Chinese tech giants, including Tencent, Alibaba, Ant Group, and ByteDance, maintain close relationships with US data labeling firms such as Surge AI, Merkor, After Query, and Turing. These US companies also serve American clients, including federal government agencies and military branches, highlighting a complex dynamic where Silicon Valley is effectively training both sides of the AI competition. The Chinese AI labs collectively spend around $500 million annually on training data from American providers.
Overall, the trade in AI training data represents a significant and somewhat overlooked aspect of the US-China AI rivalry. While chip exports to China are tightly controlled, the flow of critical training data continues largely unregulated, enabling Chinese AI developers to leverage Silicon Valley’s expertise and infrastructure. This dual supply chain underscores the challenges in managing AI technology transfer and highlights the need for a broader approach to AI-related export controls beyond hardware.