Amazon, once an online bookseller, is destroying rare books to train AI models
Amazon is reportedly digitizing and destroying rare physical books to harvest their text as training data for large language models. These hard-to-find volumes are particularly prized because AI developers have largely exhausted the pool of content freely available on the internet.
Amazon, the company that built its empire selling books before expanding into virtually every corner of commerce and technology, is now reportedly sacrificing rare physical books to feed its artificial intelligence ambitions. According to reports, the company is destroying irreplaceable volumes in order to digitize their contents for use as training data in large language models. The rationale is straightforward but troubling: AI developers are running out of fresh, high-quality text to train their systems on. The open internet has been largely scraped clean, so unique printed materials โ especially rare books not previously digitized โ represent an untapped reservoir of language data. This raises serious concerns among bibliophiles, historians, and archivists about the permanent loss of cultural artifacts in the name of technological progress. It also highlights the growing desperation among AI companies to find novel data sources as the field advances.
Amazon, the retailer that famously launched as an online bookstore in the 1990s, has come full circle in a deeply ironic and controversial way: it is now reportedly destroying rare books โ not to sell them, but to feed their text into artificial intelligence training pipelines. The company is physically digitizing and then discarding rare volumes, treating them as raw fuel for large language models rather than as cultural artifacts deserving preservation.
The underlying driver here is a crisis quietly brewing across the AI industry: data scarcity. The leading AI labs have already trained on an enormous proportion of the publicly accessible internet โ websites, forums, digitized libraries, open-access academic papers, and more. What remains untouched is increasingly found in the physical world, in printed materials that were never scanned and uploaded. Rare books, by definition, exist in limited quantities, may be out of print for centuries, and contain language, knowledge, and stylistic patterns not replicated anywhere online. That makes them extraordinarily valuable as training inputs.
Why it matters: The destruction of rare books for AI training represents a collision between two very different conceptions of value. For archivists, historians, and scholars, a rare book is an irreplaceable artifact โ its physical form, provenance, and condition carry meaning beyond the text itself. For AI developers, it is a data source to be consumed and discarded. Once destroyed, these objects cannot be recovered.
This story also illuminates a broader and uncomfortable truth about the AI boom: the industry's hunger for training data is so voracious that companies are now willing to permanently eliminate cultural heritage to satisfy it. It raises urgent questions about oversight, regulation, and ethical responsibility. Who decides which books are expendable? Are copyright holders or original creators compensated? What happens when commercially motivated data harvesting conflicts with society's interest in preserving history?
Amazon's involvement is particularly pointed given its origin story. The company once positioned itself as a democratizer of book access; it now risks being remembered as a destroyer of the very medium that gave it life. As AI development accelerates, pressure on policymakers to establish clear guidelines around physical-world data collection โ and its consequences โ will only intensify.