AI Developers Buy and Destroy Books for Training Data
Large language models need vast amounts of professionally edited, human-written text, making books published before the generative-AI boom especially valuable as training material uncontaminated by synthetic content. Anthropic bought millions of used books, removed their bindings with industrial equipment and scanned the pages for its Claude models before discarding the originals. On June 23, 2025, U.S. District Judge William Alsup ruled that converting lawfully purchased books for AI training was fair use, sharpening debate over copyright, data ethics and cultural preservation.
A 404 Media investigation reported in July 2026 that ISBNdb, a database covering more than 111 million books, was helping AI clients acquire between 1,000 and 1 million volumes per order while keeping buyers confidential. One bookseller said weekly sales jumped from no more than 20 books to several hundred after unusual orders began in April. Anthropic separately agreed in September 2025 to a $1.5 billion settlement over pirated books, a dispute distinct from the court-approved scanning of legally purchased copies.
All Coverage
1 original reportsThe Backstory
The history behind this eventNo historical echoes for this signal
Subscribe to Mark Radar Weekly
Every Friday, the week's strongest signals in your inbox. Unsubscribe anytime.