AI Companies Digitize and Destroy Books for Model Training, Sparking Copyright and Preservation Debate

AI companies, most notably Anthropic, have come under scrutiny for a controversial practice: purchasing, digitizing, and then destroying physical books—primarily out-of-print, technical, and non-fiction works—to create proprietary datasets for training artificial intelligence models. This approach, which has drawn comparisons to the burning of the Library of Alexandria, is intended to comply with copyright laws and minimize legal risks, but it has ignited a debate over cultural preservation and the future of physical books.

According to multiple reports, Anthropic and similar firms have been acquiring large quantities of books, often those that are discarded or no longer in print. After scanning the pages to create digital copies, the companies destroy the physical books. This destruction is not arbitrary; it is driven by legal considerations. A recent court ruling allowed Anthropic to digitize books for AI training if each physical copy was destroyed after scanning, thereby reducing the risk of statutory damages for copyright infringement, which can reach up to $150,000 per work.

While some viral claims allege that rare and valuable books are being lost forever, experts and investigative reports clarify that the vast majority of books targeted are not unique or rare. Instead, they are typically mass-produced volumes from the late 20th century, many of which would have otherwise been recycled or destroyed. The process is seen by some as a form of recycling, transforming forgotten paper into accessible digital knowledge.

The motivation for this practice is rooted in the need for high-quality, curated data to train advanced AI models. As the internet becomes saturated with AI-generated content, older books—especially those written before the AI era—offer valuable, structured information that is difficult to find online. This makes them ideal for training systems like Anthropic’s Claude or OpenAI’s ChatGPT.

The controversy has led to renewed calls for copyright reform and stronger support for public institutions such as libraries and digital archives. Critics argue that current copyright laws, shaped by powerful media interests, restrict access to much of humanity’s cultural heritage, especially for out-of-print works where rights holders are difficult to locate. Supporters of digitization stress the importance of preserving knowledge in accessible formats, while opponents worry about the loss of physical artifacts and the precedent set by their destruction.

As the debate continues, the practice highlights the complex intersection of technology, law, and cultural preservation in the age of artificial intelligence.

Sources

Internal sources

External sources