Skip to content

Notes ·

Amazon Is Destroying Books to Train AI. We Should Have Built the Library Instead.

404 Media tracked a shipment of rare books and discovered that it ended at an Amazon facility in Las Vegas where books are being scanned for AI training data.

The way they proved it is remarkable.

404 Media placed a tracking device inside one of the books in a large shipment and followed it across the country. It eventually arrived at Amazon's VGT3 facility in Las Vegas. Workers there told 404 Media that they receive large quantities of books, cut off their bindings so they can be scanned quickly, and destroy the physical books in the process. Amazon confirmed that it purchases books through commercial channels to help develop its products and services.

There is something painfully backwards about all of this.

We already had this argument more than 20 years ago with Google Books.

Google started working with libraries in 2004 to digitize their collections. The idea was enormous: make the contents of the world's books searchable and preserve a digital copy for participating libraries. Copyright litigation followed. Google eventually won the central fair-use case, but the larger dream of a broadly accessible universal digital library never really materialized. Today Google Books has digitized more than 40 million books, but copyrighted books are generally limited to search results, snippets and whatever access publishers permit.

We had the beginnings of a modern Library of Alexandria and somehow ended up here instead.

Now some of the richest corporations in history are buying physical books because they need clean human-written training data, cutting the books apart, scanning them into private datasets and throwing away the originals.

The law is actually helping create that incentive.

In the Anthropic copyright case, a federal judge treated training on books the company had lawfully purchased and scanned as fair use, while treating Anthropic's separately acquired library of pirated books as a different copyright problem. That does not automatically decide what Amazon can do, but the message to AI companies is fairly obvious: buying a physical copy before digitizing it puts you in a considerably better legal position.

So we have managed to create a system where destroying a legally purchased book may be the safest path to creating a private digital copy of it.

That seems like a failure.

I understand the practical argument for destructive scanning. Cutting the spine and feeding loose pages through automated scanners is dramatically easier to scale than carefully imaging a bound book. Non-destructive scanning, handling and rebinding would cost more.

But that is an optimization decision, not a law of nature.

A company with Amazon's resources could work with libraries, universities or a nonprofit preservation organization. Rare or difficult-to-replace books could be scanned nondestructively. Common books could be replaced or donated after processing when practical. Most importantly, the resulting preservation copies could become part of a long-term public archive where copyright permits.

Instead we are consuming physical libraries to produce private corporate assets.

I also don't completely accept the argument that physical books do not matter because libraries discard books all the time.

A paperback sitting in a damp garage is obviously not permanent archival storage. But paper has an enormously useful property that digital storage does not: it can remain readable while doing absolutely nothing.

It doesn't need power. It doesn't need a filesystem that somebody still understands, a functioning storage controller, an account, a subscription, a cloud provider or a company that still exists.

We can read physical manuscripts that survived centuries before anyone alive today was born. A hard drive sitting untouched for even a tiny fraction of that time is not a preservation strategy.

Digital preservation can be better, but only when someone actively maintains it, verifies it, migrates it and keeps multiple copies alive.

The physical and digital copies should complement each other, not require destroying one to create the other.

There is another uncomfortable conclusion here.

I actually want copyright law enforced against these companies.

Not because I think our current copyright system is particularly good. I don't. Congress repeatedly lengthened copyright terms, including another 20-year extension in 1998, leaving generations of culture unavailable to the public long after much of it had stopped being commercially useful. The Copyright Office itself has acknowledged how those changes created problems around unavailable and orphaned works.

But the largest technology companies have enough money and lobbying power to change laws.

If copyright remains an obstacle primarily for libraries, archivists, researchers, small publishers and ordinary people while trillion-dollar corporations find a sufficiently expensive path around it, there is very little incentive to fix the system.

Make Amazon, Meta, Google and the rest live under the same dysfunctional copyright regime and suddenly reform becomes economically interesting.

Maybe then we could have the conversation we should have finished during the Google Books era: shorter and more reasonable copyright terms, meaningful treatment of abandoned and orphaned works, strong preservation rights for libraries, and a legal path toward an open digital library of human knowledge.

AI companies clearly believe books are valuable enough to buy by the warehouse.

I agree with them.

I just wish we valued the library as much as the training data.

All notes