What We Lose When AI Eats the Public Library
For they are the destroyers of facts
Efficiency in scanning is not a sufficient reason to destroy objects that communities, libraries, collectors and future readers may value. The destruction of books for AI training troubles me in two main ways.
First, it treats cultural artefacts as disposable containers of extractable text. Secondly, it threatens the material conditions through which societies establish, preserve and contest facts. Books are more than mere information storage devices. They are stable, citable, inspectable objects that support historical memory, scientific reproducibility, and shared reference.
The couldn't-care-less destruction of books points to a messy question:
What happens to knowledge when the physical sources that stabilise it are destroyed, official records are silently withdrawn, social platforms provide different realities to different users and AI systems continuously change the way information is generated and presented?
Chronology
January 2026
The first major reporting followed the unsealing of court documents in the copyright litigation against Anthropic. The Washington Post reported in January that Anthropic had acquired millions of books, removed their bindings, scanned them and disposed of the physical copies.
An internal document described Project Panama as an effort to “destructively scan all the books in the world”. The documents reportedly described industrial processes involving hydraulic cutting equipment, high-speed scanners and recycling arrangements.
The judge found that books Anthropic had legally purchased could be used in this way under fair-use principles, while the wider copyright dispute continued. This might have answered the legal question, but it didn't answer the cultural or epistemic questions.
- The legal question concerns acquisition, copying, copyright and training.
- The cultural question concerns whether destruction is acceptable even when copying is lawful.
- The epistemic question concerns what is lost when the source object disappears.
While anyone can destroy a book (burning is a favourite of authoritarian regimes) or ban it, a problem forms when the destroyer has a stake in controlling knowledge and made worse when the book destroyed is rare.
July 2026
In July, 404 Media reported that ISBNdb had promoted a service offering printed books for AI training, with orders ranging from thousands to as many as one million titles. The company described the service as tailored to “LLM training needs” and “delivered at the scale AI demands”.
The report suggested that book acquisition was becoming an organised supply-chain activity rather than an isolated project. ISBNdb later removed the page and said it had only been testing market interest and had never purchased, scanned or sold books for AI training. That retraction illustrates how quickly commercial incentives can move ahead of public scrutiny.
August 2026
Booksellers in Australia, the UK and Ireland then report suspicious bulk orders involving rare, specialist and out-of-print titles. The Guardian described Australian booksellers’ concern that books might be entering an AI supply chain and being destroyed after scanning. In the UK and Ireland, sellers reported “scattergun” requests for large numbers of unrelated books, with one Galway bookseller describing an order for several thousand titles.
The strongest new evidence is 404 Media’s AirTag investigation. A rare book placed into a bulk shipment was tracked to an Amazon facility where workers reportedly cut bindings and scanned books. Amazon confirmed that it buys books through commercial channels to improve its products and services, although that does not establish that every book in the operation was rare or that every acquisition was specifically for model training.
This takes us to the present moment and the working conclusion:
There is clear evidence of destructive book scanning by Anthropic, and credible evidence of a broader industrial book-acquisition and scanning ecosystem involving Amazon and intermediaries. The full scale, participants and fate of rare books remain incompletely documented.
Culture and History
A book is more than text
The AI supply chain treats a book as a container:
- acquire the object;
- remove the binding;
- scan the pages;
- extract the text;
- discard the remainder.
That process preserves one dimension of the book while destroying others. It retains linguistic content while losing:
- paper, binding, typography and design;
- marginalia, inscriptions, stamps and ownership marks;
- evidence of provenance and circulation;
- physical clues about production and readership;
- the possibility of future examination using techniques not yet available.
The astronomer Tycho Brahe would attend the great European book fairs and search for second hand copies of his own work. This was less about his colossal ego and more about his curiosity about what the original buyers thought of his work: he was focused on the marginalia.
I scoured second hand bookstores for a particular type of book: it had to be written before 1939 (because up to that point no one could be certain that war would come) and its topic was the rise of fascism (because up to that point fascism was just another political movement in an era of extremes).
I collected books written by enthusiasts, opponents and enablers of fascism so I could get into their collective mind, because they were all writing into the same information environment. The most valuable are the ones with sentences underlined and observations in the margins.
This was where I could see the previous owner connect ideas and I made a few breakthroughs by following the notes. I came to understand why Brahe hunted those tiny scraps of additional knowledge and now do the same with my books, scribbling notes to connect the mess of ideas in my head.
For a common paperback or niche works (such as Born to Thunder: Champions of New Zealand Cycling, one of the books sent for destruction in the Guardian reporting), loss by destruction is modest but for a rare or unique volume, it is irreversible. A scan preserves what the operator decided to capture, we see their perspective, only their perspective and we can't see what they chose not to capture.
The analogy with burning books is reasonable, even though the motives differ. Burning destroys books to suppress or erase their contents. Destructive scanning destroys books to appropriate their contents for another system. The former is ideological destruction while the latter is instrumental destruction: both remove physical sources from the world.
The history of science matters
The history of modern science helps explain why this is more than a preservation concern.
Scientific knowledge depends on claims being inspectable, transferable, and capable of being checked by others. The printed book played a crucial role in creating that condition. Elizabeth Eisenstein described print as a technology that made knowledge mobile while preserving identical forms across copies. Bruno Latour later used the phrase “immutable mobiles” for inscriptions that can travel without losing their essential form.
David Wootton develops a related argument in The Invention of Science:
“Facts are made in the image not of people, who misremember, misquote and misrepresent, but of books, immutable and mobile. The fact is an epistemological shadow originally cast by a material reality: the printed book.”
Wootton isn't saying that words on paper are magically truthful. Books can contain errors, lies, propaganda and bad science. Their importance comes from their relative fixity; a reader can return to the same edition, a scholar can cite a page, a critic can compare copies, and a claim can be challenged against a stable source.
One outcome of the European book fairs is attention began to focus on errors, and a series of 17th century works just focused on compiling those errors. The origin of the footnote was to ensure that every fact could be traced to an authorising statement. When the books are scanned, did they get all the footnotes? How to know without comparing back to an original.
This factual stability helped facilitate a social technology of disagreement. People might interpret the same book differently but can still point to the same words. A book is one of Latour's immutable mobiles in a strong sense: it can move across time and place while remaining sufficiently stable for independent inspection. The reliability of fact is partly constituted by the stability of the book.
Destructive scanning preserves the text as data while removing the material source that anchored it. The resulting file may be searchable and reproducible within the AI company’s infrastructure, but the public loses the ability to inspect the object from which the data came. Provenance shifts from libraries, booksellers, collectors, and readers to private data pipelines.
That is a major change in the social politics of knowledge.
The ephemeral digital
Disappearing official records
The destruction of books becomes even more significant when placed alongside the instability of digital information.
In 2025, more than 8,000 pages were removed from US government websites across multiple agencies, according to The New York Times. The removals affected information about vaccines, veterans’ care, hate crimes, reproductive rights, diversity programmes and scientific research. Harvard researchers subsequently described health data and other public information disappearing from government sites, with some material later restored.
The point is not that every removed page was correct or permanently lost. The point is that a web page is often treated as if it were a stable public record when it may actually be mutable, retractable or politically contingent. I saw this during the early days of Trumf 1, when the AI strategy work by the Obama administration was taken down within days of his inauguration. Yes, I have them as hard copy because of all the notes I made and the longitudinal view they provide; they're now safely in my three-drawer filing cabinet.
A printed government report held in libraries creates a different condition. It can be removed from the agency website, but copies may remain in offices, archives, university collections and private hands. Its existence becomes difficult to erase completely. A web page can disappear for most practical purposes at the moment its host removes it.
This is the inverse of destructive scanning:
- With destructive scanning, the physical source is sacrificed to create a private digital copy.
- With web deletion, the public digital copy disappears while private or archival copies may or may not survive.
- In both cases, control over the persistence of knowledge shifts away from the public and into the hands of a few.
The distinction between a record and a live webpage becomes increasingly important. Public institutions need durable records, versioning, provenance and archival obligations. A URL alone does not an archive make.
Social media and the end of the shared page
Social media introduced another rupture. People should no longer assume they encounter the same information when they visit the same platform. Feeds are ranked, personalised and shaped by opaque systems designed around engagement and prediction.
Research on epistemic fragmentation describes how targeted digital environments can separate people from shared context. Users may not know what information others received, which makes it harder to assess competing claims and identify manipulation. Algorithmic systems can make some voices more visible, hide others and reinforce polarisation.
A book creates a common object, but a personalised feed creates a private sequence. Two people can discuss “the article” while having received different surrounding material, different headlines, different recommendations, different juxtapositions and different rebuttals. This fundamental disconnect extends beyond interpretation and into the informational environment itself.
This makes social media a different epistemic regime from print:
| Print culture | Personalised digital media |
|---|---|
| Stable artefact | Mutable presentation |
| Publicly citable source | Individually ranked feed |
| Same text available to readers | Different context for different users |
| Errors can be revisited and challenged | Content can be edited, removed, or buried |
| Provenance can be inspected | Ranking and selection often remain opaque |
AI adds a further instability
AI systems introduce a new element because they do much more than distribute information. They generate responses that can vary by model, version, prompt, user context, retrieval source and system instruction. I use AI a fair amount but am careful to put constraints around that use. As all the Big Four have now been caught using unchecked AI outputs in their (very expensive) reports, it's clear why many people distrust AI as a knowledge medium.
An AI provider can update a model without changing the product name. The same question can then produce a different answer next month. A model can receive a safety update, alter its refusal behaviour, change its source weighting or modify its retrieval process. Unless the user records the model version, system context, source material and output, the result can be difficult to reproduce.
Generative AI research has long identified reproducibility as a central challenge. Reliable evaluation requires capturing the model, data, code, settings, tools, and operating environment. Even when a model is held constant, outputs can vary because large language models are stochastic.
This changes the basis of knowledge generation:
- The source may be privately held or destroyed.
- The dataset may be inaccessible.
- The model may be updated.
- The prompt may be unavailable.
- The output may be non-deterministic.
- The response may be personalised.
- The provider may alter the system without preserving public versions.
A printed book says, in effect: this is what was written in this edition. An AI system says: this is what I generated under these conditions, which may no longer exist.
AI can help people find and synthesise knowledge: it's given me a step change in my knowledge work that would have been very hard to replicate without a team of researchers and writers. It also makes the provenance of knowledge harder to see. The more persuasive the output, the more important that provenance becomes.
The AI industry dilemma
The AI industry is consuming stable cultural artefacts to build systems that produce unstable informational experiences.
It needs books because books are dense, edited, relatively reliable, and less contaminated by synthetic text than the open web. But the systems trained on those books can still return fluid, personalised and changing answers without exposing the source pathway. Destroying the immutable mobile for an AI purpose doesn't make the knowledge output equally stable.
The industry depends on the very qualities it risks weakening:
- fixity;
- provenance;
- editorial accountability;
- cultural continuity;
- and the ability to return to an original.
AI companies are mining the material culture of stable knowledge to create a system in which knowledge is increasingly detached from stable sources.
This makes the governance of source material and model behaviour culturally important. And that's something I definitely don't trust a technologist to do, especially one from Silicon Valley.
I'll be dead in 50 years. My art will survive me and hopefully some of my printed work. But when my estate releases my books into the second-hand system, which is what happened to all the owners of the pre-1939 books I bought, I hope that someone will look at what I've underlined and scribbled in the margins. Because that points to where the good stuff is.
Unanswered questions
I haven't thought the topic through in depth yet; I'm a busy public servant up to my ears in technology change, pushing my projects along before I retire. But there are some open questions on this topic that deserve some thought because, clearly, this is a hot button topic for me and I'm writing as much with my angry lizard brain as I am with rational reasoning.
- Should rare or culturally significant books ever be destructively scanned for commercial AI training? My initial answer is no because their value belongs to many not the few.
- Should AI companies be required to disclose source acquisition and preservation practices? My initial answer is yes because we may not like what we see.
- Should libraries, archives and cultural institutions have standing in decisions involving mass digitisation and destruction? My initial answer is yes because these institutions are crucial to a functioning democracy.
- What counts as an adequate public record when government information is primarily published online? My initial answer is anywhere but a US website because the rate of webpage take downs is increasing.
- Should AI providers preserve model versions so that important outputs remain reproducible? My initial answer is yes because they were part of a knowledge generation process.
- Can an AI-generated answer be treated as knowledge without a stable chain of sources? My initial answer is yes with the caveat that it needs to find its place in that stable chain in order to contribute to and take strength from that chain.