AI Bubble: ‘AI companies have run out of internet.’ | David Gerard

David Gerard discusses how AI companies have exhausted freely available internet data for training, resorting to ethically problematic methods like destructively scanning rare books and aggressively scraping online content, which raises concerns about cultural loss and data quality degradation. He warns that the AI industry’s current unsustainable model may soon collapse, emphasizing the urgent need for better policies, responsible data sourcing, and legal reforms to protect creators and cultural heritage.

In this discussion, David Gerard highlights a significant issue facing AI companies: they have effectively “run out of internet” data to train their models. Initially, AI models relied heavily on datasets like the Common Crawl, which is a comprehensive snapshot of the internet. However, as these datasets become saturated, companies are turning to alternative sources such as bulk purchases of secondhand books, often rare or local histories, which are then destructively scanned to extract text for training. This practice, while legally supported under recent fair use rulings in the US, raises ethical concerns about the destruction of physical books and the loss of cultural heritage, especially when these books are not preserved digitally.

Gerard explains that the quality of AI outputs is directly tied to the quality and quantity of training data. As AI companies exhaust freely available online content, they face diminishing returns and the risk of “model collapse,” where training on AI-generated content leads to degraded performance. This phenomenon is likened to repeatedly saving a JPEG image until it becomes pixelated and distorted. To combat this, companies are aggressively scraping data from various sources, including websites, Twitch streams, and archived communications, often ignoring standard web protocols and causing significant strain on smaller websites.

The conversation also touches on the controversial practice of watermarking AI-generated content, as implemented by Anthropic in their Claude model. While watermarking aims to identify AI-generated text and images, it has been met with user resistance. Critics argue that watermarking does not solve deeper issues such as attribution and copyright infringement, especially in the creative arts where AI models have been trained on artists’ works without consent. Gerard emphasizes that artists harbor strong resentment towards AI’s impact on their livelihoods, and current watermarking methods fall short of providing meaningful attribution or compensation.

Looking ahead, Gerard is skeptical about the sustainability of the current AI industry model. He predicts that the AI bubble, fueled by venture capital investment rather than solid business plans, is likely to burst within the next year. This collapse could lead to a reckoning regarding the legal and ethical issues surrounding AI training data and copyright. However, he cautions that without proper legislation or industry reform, the damage to cultural and creative content may be irreversible, and the companies responsible may simply disappear without accountability.

Ultimately, the discussion underscores the urgent need for better policies and practices around AI training data. Gerard suggests potential solutions such as requiring scanned books to be archived publicly to preserve cultural knowledge and calls for more responsible data sourcing. He also highlights the broader societal implications of unchecked AI data harvesting, including the erosion of privacy and the exploitation of creators. The conversation ends on a note of uncertainty about the future but stresses the importance of public awareness and legislative action to address these challenges.