The video debunks viral claims that AI companies like Anthropic are recklessly destroying rare books, explaining that they primarily digitize discarded or out-of-print books and destroy physical copies to comply with copyright laws and reduce legal risks. It emphasizes the need for copyright reform and stronger support for public institutions like libraries and the Internet Archive to preserve cultural knowledge amid legal and economic challenges.
The video addresses recent viral claims that AI companies, particularly Anthropic, are buying massive quantities of books, scanning them for AI training, and then destroying the physical copies. Many people have expressed outrage, fearing that rare and valuable books are being lost forever, likening the situation to burning the Library of Alexandria. However, the reality is more nuanced. The books Anthropic acquires are generally not rare or valuable but rather discarded or out-of-print books that would likely have been destroyed or recycled otherwise. This bulk acquisition is a cost-effective way for AI companies to gather large datasets for training their models.
Anthropic is not the first company to digitize large volumes of books. The Google Books Project, started over 20 years ago, digitized millions of books in partnership with libraries and publishers, but crucially, Google returned the physical copies to the libraries rather than destroying them. Google’s project aimed to make books searchable and accessible, and although the product itself is less known today, the digitized corpus remains a significant asset for AI development. Anthropic, trying to compete, has adopted a different approach, buying books in bulk and scanning them, but the destruction of physical copies has sparked controversy.
The key reason Anthropic destroys the books after scanning is rooted in copyright law and legal risk. Copyright infringement can lead to statutory damages of up to $150,000 per work, which can be financially devastating when dealing with millions of books. A court ruling allowed Anthropic to digitize books if they made a one-to-one replacement by destroying the physical copy, which reduces legal liability. This legal framework incentivizes companies to destroy physical books after digitization to avoid massive copyright infringement penalties, despite the emotional and cultural discomfort this practice causes.
The video also highlights the broader challenges posed by copyright law, especially for out-of-print books where rights holders are difficult or impossible to locate. Licensing is often not a feasible solution for digitizing and preserving many books, particularly those from the 20th century. The current copyright system, shaped by powerful media companies, restricts access to much of our cultural heritage, creating a deadweight loss where valuable knowledge remains locked away. Reforming copyright law and supporting public institutions like libraries and the Internet Archive are suggested as necessary steps to preserve knowledge for future generations.
Finally, the video stresses that while AI companies like Anthropic are caught in a difficult legal and economic bind, the responsibility for preserving human knowledge should not rest solely on them. Public institutions and archives play a crucial role in this mission, but they too face legal challenges and funding shortages. The Internet Archive, for example, is under legal attack despite its importance for journalists, historians, and the public. The video calls for greater public support for these institutions and copyright reform to ensure that knowledge is preserved and accessible, rather than lost or locked away due to outdated laws and commercial pressures.