Key Takeaways

  • The operation: Amazon buys printed books in bulk, strips their bindings, scans the pages, and destroys the volumes to train its own artificial intelligence models.
  • The discovery: An investigation by 404 Media, published on August 17, 2026, used an Apple AirTag to track a batch of 1,000 books to Amazon's VGT3 operations center in Las Vegas.
  • The context: The practice, already adopted by Anthropic under "Project Panama," targets pre-2022 texts to avoid "model collapse" and has run into lawsuits centered on fair use.

The Investigation and the Book's Trail

An investigation by 404 Media revealed that Amazon has spent at least two years running a large-scale operation to buy printed books with the sole purpose of digitizing and destroying them. Suspicions first arose from independent booksellers who flagged unusual bulk orders of rare, out-of-print titles placed by buyers indifferent to price.

To verify the lead, journalists worked with a bookseller who had received an order for roughly 1,000 books through the Biblio platform. An Apple AirTag hidden inside a rare volume tracked the shipment's journey: from a California airport to Milwaukee, then by truck to a distribution center in Wisconsin, and finally to Amazon's VGT3 warehouse in Las Vegas, Nevada.



Amazon Destroys Print Books to Train Its AI - Foto 1

Employees at the facility confirmed receiving "massive shipments of printed books," cutting the bindings to speed up page scanning, and then destroying the volumes. The internal logo of the VGT3 team, a dinosaur clenching a book in its jaws, sums up the operation neatly.

Why Print Books Matter for Training

Large language models require growing volumes of quality text, and AI companies have already exhausted much of what's available online. Books printed before 2022 hold a specific value: they contain text not digitized anywhere else and, crucially, were written before generative AI became widespread.



Amazon Destroys Print Books to Train Its AI - Foto 2

This second point is decisive. Training a model on text generated by other AIs triggers "model collapse," a progressive deterioration in response quality. Pre-2022 books therefore represent a source of text "clean" of algorithmic contamination.

The Legal Front and Prior Cases

Amazon isn't alone in this. Anthropic, through its so-called "Project Panama," ran similar operations buying and destroying millions of books to train its Claude chatbot. In June 2025, a US federal judge ruled that buying books, dismantling their bindings, and scanning them can fall under fair use, effectively clearing the way for these practices.

The line separating this from copyright infringement remains thin, though. Anthropic has already settled a class-action lawsuit from authors who challenged the unauthorized use of their work, paying out $1.5 billion. Amazon has only stated that it purchases books "through commercial channels to help develop and improve the products and services" it offers customers.

How the Titles Are Selected

Booksellers interviewed noticed that Amazon's orders systematically exclude volumes without an ISBN, a detail suggesting a methodical goal: scanning every printed book with standard cataloging. But destroying rare, out-of-print editions causes a loss that goes beyond text content, touching the historical and material value of the objects themselves.



Amazon Destroys Print Books to Train Its AI - Foto 3

The company that launched online book sales in 1995 now buys books only to physically eliminate them after extracting their content. The case fuels a debate set to intensify as AI companies keep searching for data sources untainted by artificially generated content.