AI Firms Quietly Buy and Destroy Books for Training Data
Generative AI developers need vast volumes of polished, human-written text to train language models, making books published before the 2022 boom in synthetic content especially valuable. Anthropic’s Project Panama established that at least one major lab bought and destructively scanned physical books for Claude, cutting off bindings before digitizing the pages and discarding the originals. The practice has sharpened disputes over copyright licensing, the loss of rare cultural material and the transfer of publicly available knowledge into proprietary datasets.
404 Media reported on July 21, 2026, that ISBNdb was offering anonymous AI clients bulk sourcing of physical books, with orders of up to 1 million volumes drawn from used-book stores and out-of-print catalogs. Anthropic separately began its industrial scanning effort in February 2024, spending millions of dollars to process millions of books. A U.S. federal judge ruled in June 2025 that converting lawfully purchased copies and using them for model training was fair use, while distinguishing that practice from the company’s acquisition of pirated digital books.
All Coverage
2 original reportsThe Backstory
The history behind this eventNo historical echoes for this signal
Subscribe to Mark Radar Weekly
Every Friday, the week's strongest signals in your inbox. Unsubscribe anytime.