• ISBNdb ships up to a million books anonymously to AI labs
  • Pre-2022 books are prized because chatbots never touched their text
  • Spine-cutting scanners destroy originals to speed up digitization for training

AI companies are increasingly turning to printed books published before 2022 as preferred training material because those works predate the widespread use of AI-generated content.

Large-scale scanning operations reportedly involve cutting book spines, separating pages, and destroying physical copies to create digital datasets for large language models.

The practice has attracted growing criticism because some books entering these pipelines are reportedly extremely rare, raising concerns about irreversible cultural losses.

Pre-2022 books become valuable AI training material

Reports by 404 Media found data broker ISBNdb supplies physical books in bulk to AI developers seeking human-written material unaffected by modern chatbot output.

The company argues books published before 2022 offer cleaner datasets because they cannot contain text generated by contemporary large language models.

They are often considered "dense, edited, authoritative," in contrast to internet content increasingly filled with machine-generated material of uncertain quality.

The approach also attempts to avoid so-called model collapse, in which AI systems gradually lose quality after repeatedly training on synthetic content generated by earlier models.

ISBNdb additionally argues that older printed works avoid deliberate data-poisoning techniques authors increasingly use to disrupt AI training through carefully modified documents.

However, there are reports that many of these books are scanned using high-speed equipment.

This equipment requires workers to remove the spine before feeding individual pages through automated imaging machines.

That process reportedly destroys the original volume, making rapid digitisation considerably cheaper than slower preservation methods designed to keep books physically intact.

Secrecy and legal rulings fuel preservation concerns

ISBNdb openly acknowledges reputational concerns surrounding the practice while offering strict non-disclosure agreements that keep customer identities confidential throughout commercial engagements.

Its website reportedly states, "'AI company destroys two million books' is not a headline that generates sympathy," while suggesting clients describe the process as digital preservation.

Such a level of destruction is an order of magnitude bigger than the loss of the Library of Alexandria. Yet, it is unfolding with none of the outrage that history reserves for burned libraries.

Booksellers interviewed by 404 Media said some volumes entering these scanning programmes have very few surviving copies after enduring wars, fires, and centuries of handling.

Critics argue that unlike websites or widely available modern publications, exceptionally scarce historical works cannot simply be reproduced after their physical copies disappear forever.

A recent United States court ruling involving Anthropic found that scanning legally purchased books for AI training constituted fair use under specific circumstances.

Part of that reasoning held that destroying each printed copy during scanning meant one legal copy effectively replaced another rather than creating multiple copies.

In response to a critic (@Hedgie) of this method on X, Elon Musk said, "I've asked the SpaceXAI team to preserve any rare books in a library and scan them the hard way," suggesting an alternative approach.

If significant awareness is not created, this quiet erasure of irreplaceable books risks becoming the defining act of cultural loss for this era, remembered only after it can no longer be undone.

Efosa has been writing about technology for over 7 years, initially driven by curiosity but now fueled by a strong passion for the field. He holds both a Master's and a PhD in sciences, which provided him with a solid foundation in analytical thinking.