Notes ·
Aaron Swartz Was Prosecuted. Meta Torrented Shadow Libraries at Industrial Scale
A lot of people know Aaron Swartz as an internet activist. Fewer seem to know just how much of the early web he helped build.
At 14, Swartz was part of the RSS-DEV Working Group that authored RSS 1.0. His company Infogami later merged with Reddit, where he became an equal owner of the resulting company and helped rewrite Reddit using Python and web.py.
Then there is what happened with JSTOR.
Swartz downloaded more than 4 million academic articles through MIT's network. The federal government charged him with computer and wire fraud offenses and said he faced as much as 35 years in prison and a $1 million fine.
What is particularly difficult to reconcile is that JSTOR had already settled its civil dispute with Swartz. It recovered the downloaded material and said it had no interest in the matter continuing. The federal prosecution continued anyway.
Swartz died by suicide in 2013 while the case was still pending.
Now compare that with Meta.
There is no dispute in the current litigation that Meta torrented material from LibGen and Anna's Archive while gathering training data for Llama. A 2026 lawsuit from major publishers goes further, alleging that Meta downloaded 134.6 terabytes through torrents during one period in 2024 and uploaded more than 40 terabytes back to other peers.
Meta denies wrongdoing and argues that using copyrighted material to train AI can constitute fair use.
These cases are not legally identical. Swartz circumvented attempts to block his downloads and was prosecuted under computer crime and fraud laws. Meta is fighting copyright claims involving both the acquisition of the material and how it was subsequently used for AI training.
But the difference in proportionality is difficult to ignore.
An individual downloading an academic archive in pursuit of broader access to knowledge faced federal prosecutors threatening decades in prison. One of the largest corporations in the world can acquire enormous collections from known shadow libraries for a commercial AI system, and the question is largely being worked out through civil litigation where money is the likely consequence.
I don't think that says much about the difference between downloading four million articles and downloading millions of books.
It says something about the difference between doing it as Aaron Swartz and doing it as Meta.