AI companies cut spines off millions of books to feed their models

AI companies are buying millions of books, slicing off spines, and scanning pages before recycling them for training data. ISBNdb abandoned its book-destruction service after a backlash, yet Anthropic already spent millions to destroy up to 2 million books in six months for copyright settlement.

Categorized in: AI News Writers
Published on: Aug 09, 2026
AI companies cut spines off millions of books to feed their models

The AI industry's appetite for fresh training data has led some companies to buy millions of physical books, slice off their spines, and scan the pages before recycling the remains. A backlash last week forced ISBNdb, a book-data company, to abandon its offer of exactly that service - and the episode shows how far AI companies will go to train their models on well-edited text.

ISBNdb, which tracks book data, offered "AI training data" in the form of "the world's largest book database," according to a 404 Media report. A now-deleted page also advertised a "legally-binding nondisclosure agreement" that would hide an AI client's "identity and strategy." The company explained why in blunt terms: "Destroying millions of books evokes images of burning libraries ... the optics problem is real. 'AI company destroys two million books' is not a headline that generates sympathy."

The service worked by slicing spines off books and running pages through a scanning machine, photocopier style, before recycling them. ISBNdb's now-deleted advice to clients: "lead with the recycling story, not the destruction story."

The destruction story caught the internet's attention anyway. A financial newsletter post about the practice drew 2,000 outraged replies, including one from Elon Musk, who pledged that Grok's book training would be done the "hard way" in the case of rare books. Commenters pointed out that Musk offered no definition of "rare book" - nor any promise his company would avoid destroying books in other cases.

ISBNdb has since walked back the offering. "We've chosen to pivot away from that direction," the company's news site says, calling the service a "test of market interest." ISBNdb assured users it hadn't scanned a single book.

Books are attractive training material for a simple reason: they are generally better edited and more cogent than the average internet thread. There is also a premium on books published before 2022, because no one can be sure whether books written after that date were produced by AI.

Anthropic's Project Panama

ISBNdb is not alone. Anthropic admitted to using millions of pirated books to train Claude, and a class-action lawsuit that ended in a record $1.5 billion copyright settlement revealed the scale of its physical book-destruction operation.

U.S. District Judge William Alsup described the effort, called Project Panama, in a June 2025 legal order. Anthropic "spent many millions of dollars to purchase millions of print books, often in used condition. Its service providers stripped the books from their bindings, cut their pages to size, and scanned the books into digital form - discarding the paper originals."

Anthropic sought a vendor to "convert from 500,000 to two million books over a six-month period." One vendor boasted of its "hydraulic powered cutting machine" and said a recycling company would pick up the remnants.

Alsup ruled the destruction fell under fair use, since the company was effectively exchanging one real-world copy for one digital copy. As a matter of law, that protected Anthropic. As a matter of public relations, it generated exactly the headlines ISBNdb wanted to avoid.

The slower, gentler alternative

Nondestructive scanning exists. The Internet Archive has been scanning books for years with its Scribe system, where every page is turned by hand. One operator said she had scanned roughly 18,000 books in 10 years.

The Archive uploaded its 2 millionth book in 2021, after 20 years of work. That is roughly the pace of one year's new book production, which stood at about 2 million titles annually before 2022. Anthropic's vendors aimed to scan that many in six months.

There is no sign AI companies will slow down on their own. The thing that made ISBNdb back down was public scrutiny - the same spotlight that turned Anthropic's court victory into a public relations disaster.

Why this matters for writers

Your books are the raw material for this industry. Companies training large language models want well-edited, carefully structured text, and they have shown they will buy up physical copies and destroy them to get it. The $1.5 billion settlement in the Anthropic case shows authors can win compensation after the fact, but it does nothing to stop the next project. Public attention - the kind that forced ISBNdb to retreat and pushed Musk to promise a "hard way" for rare books - is the only check that has worked so far.


Get Daily AI News

Your membership also unlocks:

700+ AI Courses
700+ Certifications
Personalized AI Learning Plan
6500+ AI Tools (no Ads)
Daily AI News by job industry (no Ads)