Data Quality Is the Compute Multiplier — Ari Morcos, DatologyAI

Ari Morcos, CEO of Datology AI, emphasizes that improving data quality serves as a powerful “compute multiplier,” enabling machine learning models to achieve better performance with less compute by maximizing the marginal information gain per data point through careful curation and refinement of datasets. He presents evidence from various domains showing that curated data leads to more efficient and cost-effective model training, advocating for a shift in focus from merely increasing data volume to enhancing data signal quality, especially in compute-constrained environments.

Ari Morcos, CEO and co-founder of Datology AI, opens his talk by emphasizing the critical importance of data quality in the current landscape of machine learning, especially as compute resources become increasingly scarce and expensive. He highlights that while compute availability has tightened, with hardware costs rising and token usage skyrocketing—particularly in reasoning models—improving data quality acts as a powerful “compute multiplier.” By enhancing data quality, models can achieve significantly better performance with the same or even less compute, effectively bending traditional scaling laws and making the learning process more efficient.

Morcos explains that maximizing the marginal information gain per data point is key to improving data quality. This involves curating data that is highly relevant to the specific tasks the model is intended to perform, ensuring diversity to avoid brittleness, and carefully mixing data sources. Datology AI approaches this challenge by refining existing datasets rather than sourcing new tokens, using a four-step process: clean, curate, create, and compose. This includes rigorous cleaning, benchmark decontamination to avoid data leakage, quality classification, redundancy reduction, and synthetic data generation through rephrasing to increase both size and diversity of datasets.

The talk presents compelling evidence of the impact of data curation on model performance across different domains. For vision-language models, Datology’s curated datasets enable models to outperform leading public models while using significantly less compute. Similarly, in multilingual text models, careful curation leads to better performance even with limited non-English data, and interestingly, improvements in English data curation also benefit non-English language performance due to cross-lingual transfer effects. These results underscore the broad applicability and effectiveness of data quality improvements.

Morcos also discusses the practical benefits observed by Datology’s customers. For example, Thomson Reuters enhanced their legal models by mid-training on a curated combination of proprietary and public data, achieving notable gains in legal reasoning without sacrificing general capabilities. This improved data quality also amplified the effectiveness of subsequent post-training steps. Another customer, RCI, successfully trained a competitive large-scale model on 17 trillion curated tokens from public datasets for under $20 million, demonstrating that high-performance models can be built cost-effectively with the right data strategy.

In conclusion, Morcos stresses that data quality remains the most underutilized lever for improving model performance, especially in a compute-constrained environment. He advocates for focusing on obtaining the highest signal per token rather than simply increasing data volume, and highlights the need for advanced research and engineering to score and manage data at massive scales. Datology AI is actively hiring to further this frontier, inviting those interested in building or customizing models with superior data to join their efforts.