Uncategorized

AI companies keep destroying old books. Heres why.

Old books being scanned by human employees in the Internet Archive

The AI industry, not content with alienating top creators, seems to be intent on alienating book lovers too.

That much is clear in the wake of a 404 Media report on ISBNdb, a company that tracks book data and briefly offered “AI training data” in the form of “the world’s largest book database.”

So far so normal; AI companies looking to build better AI models are hungry for fresh data from any source that isn’t the internet. Books — generally better edited and more cogent than your average Reddit thread — make AI sound smart. (There’s a premium on books published before 2022, ironically because we can’t be sure if books were written by AI after that date.)

But on a now-deleted page, ISBNdb also advertised a “legally-binding nondisclosure agreement” that would hide an AI company client’s “identity and strategy.” Why? ISBNdb explained: “Destroying millions of books evokes images of burning libraries the optics problem is real. ‘AI company destroys two million books’ is not a headline that generates sympathy.”

Indeed. Turns out ISBNdb was offering destructive book scanning, an industrial process that claims to turn pages into data faster than the “move a page, take a picture” alternative — by slicing the spines off books and feeding the pages into a scanning machine, photocopier style, then recycliny them. (ISBNdb’s now-deleted advice: “lead with the recycling story, not the destruction story.”

But if you just recoiled with horror at the destruction story, you’re not alone. It wasn’t a headline so much as an impassioned pro-book post from a financial newsletter that caught the internet’s attention last week.

Among the post’s 2,000 outraged replies was one from Elon Musk, pledging that Grok’s book training would be done the “hard way” in the case of rare books.

Though as commenters quickly pointed out, Musk offered no definition of what was a “rare book” — nor any promise that his AI company would avoid destroying books in other cases.

As for ISBNdb, the company has belatedly walked back its destructive book offerings. “We’ve chosen to pivot away from that direction,” the company’s news site now says. The service was a “test of market interest.” ISBNdb assured users it hadn’t so much as scanned a single book.

But ISBNdb isn’t alone. AI-feeding book destruction has happened before, and if the rare books sales spikes also noted in the 404 are to be believed, it may be happening again, right now, with other vendors under other nondisclosure agreements.

All while other less-destructive book scanning options — with less risk of messing up the scanned text — are widely available.

Which AI companies are destroying books?

It has long made sense to train Large Language Models on books. At first, that meant books legally online or in the public domain, such as the contents of Project Gutenberg. But rivalry between OpenAI, Anthropic, and other LLM builders led to a growing temptation to use online collections of pirated books.

We may never know whether OpenAI gave in to that temptation. In 2024, a class-action lawsuit from the Authors Guild found that the company had deleted the database of books it used to train its GPT models, and no one who worked on the database still worked at OpenAI.

But Anthropic was caught, and has admitted to using millions of pirated books to train Claude. That class action lawsuit, in which a record $1.5 billion copyright settlement was just approved, also revealed the scale of Anthropic’s physical book-destruction operation.

Anthropic “became convinced that using books was the most cost-effective means to achieve a world-class LLM,” wrote U.S. District Judge William Alsup in a lengthy legal order dated June 2025.

The company was “not so gung-ho” about using pirated material from 2024 onwards, to quote a memorable internal email in evidence, but still wanted to train Claude on “all the books in the world.”

The result was called Project Panama. Alsup explains: “Anthropic spent many millions of dollars to purchase millions of print books, often in used condition. Its service providers stripped the books from their bindings, cut their pages to size, and scanned the books into digital form — discarding the paper originals.”

As a matter of law, Alsup found this book mutilation was “transformative” and fell under the definition of fair use — because of the IRL book destruction, they were effectively exchanging one real-world copy for one digital copy. If Anthropic had ignored pirated databases, bought physical books from the start, then torn them apart in this manner, the authors wouldn’t have a case.

But as a PR matter, Alsup’s order was a disaster for Anthropic — generating exactly the sort of “AI company destroys millions of books” headlines that ISBNdb was offering to reframe. In court documents, one vendor boasted of its “hydraulic powered cutting machine” and how a recycling company would pick up the remnants.

Is there a better way to scan books?

The scale and speed of Project Panama, by all accounts, were astonishing. Other court documents say Anthropic sought a vendor to “convert from 500,000 to two million books over a six-month period.”

To put that in perspective, around 2 million new books were published each year (again, before the 2022 AI slop book cutoff point). Scanning a year’s worth of books via six months of slicing and dicing is a pace at which nondestructive book-scanning methods might struggle to keep up.

Nevertheless, in a world where we’re still haunted by Nazi book burnings, as well as by classic dystopian novels of book destruction (Ray Bradbury’s Fahrenheit 451, George Orwell’s Nineteen Eighty-Four), it is the more sensitive nondestructive scanners that maintain the moral high ground.

Take the Internet Archive, a nonprofit that preserves the internet (including those deleted ISBNdb pages) via the Wayback Machine. The archive has been scanning books for years, and in 2021, a tweet about its nondestructive Scribe scanner, above, went viral — in a good way.

The Archive has experimented with scanners where robots turn pages, but found it didn’t work for rare books. Instead, every page is turned by hand. The operator in the viral tweet, one of 70 Scribe workers, said she had scanned roughly 18,000 books in 10 years.

As the Scribe was going viral, the Archive celebrated uploading its 2 millionth book — but it had taken 20 years to get there (and would soon face a copyright lawsuit of its own). Sadly, there’s no way it can keep pace with a book-slicing automated document feeder.

Instead, it’s possible the only thing that might keep AI companies from chopping up books at a rate that would make Bradbury and Orwell spin in their graves is the same thing that made IBSNdb back down: the spotlight of public attention.

Show More

Related Articles

Leave a Reply

Your email address will not be published. Required fields are marked *

Back to top button

Adblock Detected

Please consider supporting us by disabling your ad blocker