This Is Troubling: Tech Companies Are Accused of Destroying Large Quantities of Books, Including Rare/Unique Titles

“This isn’t just a copyright issue, it’s a free speech, social justice, and cultural heritage issue.” – Copyright Alliance

Reports of two tech companies destroying mass quantities of books, including rare titles, after they have been scanned for LLM training purposes are in the news.

According to an investigation by 404 Media, “Amazon is buying massive quantities of books, scanning them for AI training data, and destroying them in the process.” They placed a tracking device in a rare book that ultimately ended up at Amazon’s VGT3 warehouse in North Las Vegas where books are scanned and destroyed. In an August 26 post, 404 Media interviewed an employee who worked at the warehouse and witnessed spines being cut off books via a machine before being scanned. Afterward, the pages were “thrown into a big shuttle” making it impossible to put the books back together. The type of books scanned were varied including new ones, discards from libraries, and non-English titles. Outside of Amazon’s VGT3 warehouse there is signage of a dinosaur (with big teeth) holding a book in its claws that appears to be ready to shred or devour the book.

In January 2026, The Washington Post reported that Anthropic had a secret project, with the codename “Project Panama” to “destructively scan all the books in the world.” This secret plan was discovered in documents that were unsealed in legal filings involving its case against authors whose works had been pirated for LLM training. In these internal documents, Anthropic does define Project Panama as their “effort to destructively scan all the books in the world.” They use a “soft codename” because they “don’t want it to be known that we are working on this.” In discussing their plans for buyable new data they state, “most of our data is trained on web crawled data, and there is a limit on how much and in what categories/useful capacities data can be crawled online…beyond the data that is widely available on the web, the highest volume of useful data is available to us via published books.” “Unique, differentiated data could give us a competitive edge for others who don’t have access to the same data.” Concerning what they buy they write, “we care most about sheer volume of books rather than specialization at this time. Though, penetrating our SOM may require us to tap more into these specialized categories that are harder to find and higher priced.”

Implications

A quote in an opinion piece in The Herald identifies the central core of the issue: “Once a final non-digitized copy of a book is shredded, it’s gone forever…And the version that is digitized becomes at the mercy of AI interpretation rather than a straight, accurate representation of the work.”

To this point,Victoria Livingstone, in her thoughtful Substack article, provides examples from her own research to illustrate how authors and editors made small adjustments between printings and these changes are very often not indicated in any paratext. Even minor edits in various editions provide important insights.

The Copyright Alliance notes this “isn’t just a copyright issue, it’s a free speech, social justice, and cultural heritage issue” and losing access to original sources is especially concerning as AI-generated hallucinations are still a major problem.

Image by Michal Jarmoluk from Pixabay

September 1, 2026

Copyright © Copyright 2024 Cottrill Research. Site By Hunter.Marketing