What Evidence Actually Shows About Destroyed Rare Books
This piece examines the claim that AI companies destroying rare books. It parses booksellers' reports, what court documents confirm about Anthropic, and why blaming Amazon outright is not supported by available evidence.

Short answer: the claim that "AI companies destroying rare books" describes a mix of confirmed behaviour and marketplace suspicion. Court documents and reporting confirm that Anthropic bought and destructively scanned large numbers of books; booksellers report unusual bulk purchases through marketplaces; claims that Amazon itself is systematically destroying rare books remain contested and are reported as investigative claims, not settled fact. [S2] [S1] [S3]

Credit: Photo by Pexels on Pixabay
Why This Matters
The headlines feel urgent because physical books are cultural objects and their destruction is irreversible. When institutions or private collectors lose unique copies, researchers and future readers lose sources that cannot be reproduced by asking an LLM for a summary. That makes the mechanics of how books become digital training data, and who is buying them, a policy and preservation question as much as a tech one.
AI companies destroying rare books: what booksellers reported
Booksellers in several countries reported unusual bulk orders for obscure or out-of-print titles and suspected those purchases were for AI-training purposes. In some cases sellers initially thought the requests were spam because the quantities and selection looked random by normal resale standards. [S1]
These anecdotes share a pattern: packages with many hard-to-find titles, payments that clear quickly, and buyers who do not ask about provenance, condition beyond 'good', or resale margins. Sellers described transactions routed through online marketplaces as well as direct commercial channels (marketplaces cited in reporting include historically used platforms). Reporters and booksellers are careful to frame these as suspicions about demand, not as a forensic chain proving which company bought which box.
(Yes, watching a pallet of 19th-century monographs vanish into a shipping label labelled "data acquisition" is the sort of detail that makes cataloguers speak softly for days.)
What is confirmed about Anthropic
court filings and investigative reporting confirm Anthropic purchased and destructively scanned large quantities of physical books as part of a documented project. Documents unsealed in litigation describe practices that included buying books in bulk, slicing bindings to feed pages into high-speed scanners, and disposing of the remains after digitisation. [S2]
Those documents formed the basis for reporting that the company planned extensive scanning under an internal project name and contracted vendors capable of high-volume digitisation. The presence of legal filings means this part of the story is not mere rumour; it is grounded in documents introduced in court. That does not resolve every legal or ethical question about the practice, but it does move the claim from anecdote into documented corporate activity. [S2]
How destructive scanning works
destructive scanning is a production technique where books are prepared to be scanned at industrial speed by removing bindings and running loose pages through high-speed scanners, after which the physical remainder is typically recycled or pulped. The process trades physical integrity for throughput. [S2]
In practical terms this means a vendor can convert thousands of pages an hour when bindings are removed. Libraries and preservationists normally avoid this method for unique or rare items because it destroys the object. Institutions that do destructive digitisation tend to do so only when a copy is readily replaceable or when there is no prospect of preserving the original in situ.
Practically speaking, destructive scanning is attractive for groups that need very large, clean corpora of text at scale. That explains why some companies and vendors pursued it: it is efficient and yields consistent, page-level images for OCR. The trade-off is obvious and permanent-pages exist as data, not as a book anymore. [S2]
AbeBooks vs Amazon: why the marketplace distinction matters
reports of anonymous bulk orders ran through marketplaces including listings on platforms tied to larger corporations, but marketplace listings alone do not prove who ultimately bought or processed the books. Some marketplaces are owned by larger firms; transactions on those platforms can be resold, routed, and aggregated in ways that complicate attribution. [S1]
That matters because several stories blurred two separate claims: that orders appeared on an Amazon-owned marketplace, and that Amazon corporate teams were purchasing and destructively scanning rare books. Marketplaces can be both reselling venues and commercial conduits for third-party buyers. Where a listing comes from is not the same as who opened the box and ran it through a scanner.
Investigative reporting that tracked a shipment to what the outlet identified as an Amazon facility reported that facility performed scanning and disposal activities. That is a claim by that outlet; other reporting and sources call for caution before turning such reportage into a definitive accusation that Amazon corporate policy is to target rare books for destruction. Attribute carefully: marketplace route does not equal corporate end-user unless corroborated. [S3] [S1]
Copyright and fair-use context
the legal status of using purchased physical books for model training sits at a mix of copyright, contract, and litigation outcomes, and differs by jurisdiction. Litigation and judicial rulings have started to address whether using books to train models is infringing, but the legal landscape remains contested and fact-specific.
What matters for readers and sellers is twofold. One, a court ruling in one case does not automatically settle the ethical or policy debate for every company or country. Two, decisions about buying physical books, digitising them, and using the resulting data are not purely technical; they are shaped by contracts with vendors, the terms of sale on marketplaces, and evolving interpretations of fair use or fair dealing. Reporting on corporate plans and court documents provides evidence for behaviour; legal outcomes provide a different kind of resolution that may take years.
Cultural preservation concerns
destructive scanning creates permanent losses when unique or rare copies are consumed, and the preservation community is right to treat that outcome as a serious cultural risk. Libraries, archives, and scholars rely on physical artifacts for provenance, annotations, marginalia, and the material history that models cannot capture.
The worry is also practical. Many rare or out-of-print works exist only in particular collections or private holdings. If those copies are consigned to pulp after scanning, the original research object is gone. Even if the text survives as data, metadata about binding, ownership marks, and physical form may be lost. That is why many preservationists argue for careful triage: decide case-by-case which copies are suitable for destructive digitisation and which must be preserved intact.
What most coverage misses
broad headlines often compress distinct claims into a single narrative, which makes provenance and responsibility unclear. Coverage that shifts from "booksellers see bulk orders" to "Amazon is destroying rare books" skips necessary steps of attribution and evidence. [S1] [S3]
Two additional gaps show up regularly. One, reporting sometimes treats marketplace listings as direct proof of corporate action rather than a clue needing follow-up. Two, the preservation and legal angles rarely get the same attention as the investigatory drama (tracking a shipment makes for good headlines; court filings and copyright nuance do not).
That pattern matters because good decisions-policy, purchasing, or career moves-depend on separating what is confirmed from what is alleged. If you are watching this topic as a reader deciding whether to be outraged, to lobby, or to pursue work in data acquisitions or digitisation, focus on the documents and the supply chain rather than the most clickable headline.
Frequently Asked Questions
Why are AI companies buying rare books for training?
Many AI teams seek diverse, high-quality textual sources that are not available online. Obscure, out-of-print, or specialised books can provide material that fills gaps in web-derived corpora. Buyers targeting such texts often want content that improves model breadth and factual depth, especially in niche domains. Reporting and court documents show this motive is part of why bulk purchases have been made. [S2]
Are AI companies really destroying books after scanning them?
Some documented projects used destructive scanning for efficiency, and court documents describe this practice in at least one major case. But whether every bulk order reported by booksellers led to destruction is not settled across the market. Sellers report suspicious orders; some investigative reporting traces shipments and facilities; court filings confirm destructive scanning in specific vendor contracts. Attribute each claim to its source. [S1] [S2] [S3]
Did Anthropic destroy books to train Claude?
Court documents and investigative reporting confirm Anthropic purchased and destructively scanned books under an internal project name; those documents were unsealed in litigation and reported on by multiple outlets. That is a documented, source-backed claim rather than mere marketplace suspicion. [S2]
Is Amazon destroying rare books for AI training?
Some investigative reporting tracked shipments that, according to that reporting, arrived at a facility operated by Amazon and described scanning activity there. That reporting is an important piece of evidence but remains an investigative claim; other outlets and documents call for caution before treating it as an established corporate policy. Marketplace listings tied to Amazon-owned platforms appear in reports, but marketplace route is not the same as confirmed corporate action. Attribute carefully. [S3] [S1]
Why are booksellers seeing unusual bulk orders from AI buyers?
Booksellers report orders that deviate from normal resale patterns: large quantities of obscure titles, minimal negotiation over condition, and fast payment. Those patterns are what prompted suspicion, and they appear consistently across multiple anecdotes reported in the press. Sellers interpret these as likely for data acquisition rather than resale. Reporters have used those patterns to trace possible supply chains and buyers. [S1]
Final Thoughts
Most readers-and many headlines-collapse a handful of documented facts, investigative leads, and seller suspicions into a single, outraged claim that a named giant is buying and destroying cultural heritage at scale. That leap is understandable, but it often blurs provenance and evidence.
A more useful move is specific: read the court documents where available, treat bookseller reports as primary clues, and map the supply chain before naming a company as the actor. If you care about preservation and policy, press for transparency around vendor contracts, triage rules, and chain-of-custody for digitisation projects.
It is not an easy or fast process; the paperwork is boring and the tracking is tedious. But if you want to influence the outcome, replace viral certainty with a checklist: document, attribute, and escalate to preservation bodies or legal counsel where necessary.
If you are thinking about how this affects careers, start by mapping which organisations acquire and process physical corpora and what roles exist there-data acquisition, vendor management, legal compliance, and digitisation engineering-to discover where your skills might make an impact.
Keep reading
Related guides picked for this topic.

Groq Recasts from Chipmaker to Inference Neocloud Now
Groq raises 350 million in a Series A that values the company at $3.5B and follows a June $650M growth round. This piece explains the two raises, the neocloud pivot, and what candidates should look for when assessing AI infrastructure roles.

What the Grok Deepfake Lawsuit Means for Image Safety
The Grok deepfake lawsuit examines allegations that xAI’s Grok produced nonconsensual sexual images, including claims involving teenagers. This insight parses the separate complaints, what they allege, and the safeguards image generators must add to reduce harm.

Dario Amodei on the AI Trust Crisis: What Actually Helps
An evidence-first look at the Dario Amodei AI trust crisis. Amodei argues public scepticism comes from broader corporate credibility gaps; this piece explains why tangible outcomes, not PR, are necessary and what counts as convincing evidence.
More from AllyNerds
Not directly related — other guides readers find useful.

Rejected After Final Round Tech Interview? ,What Went Wrong?
Getting rejected after a final-round tech interview is brutal. Learn the hidden reasons why companies pass on strong candidates and how to recover

Recruiter Ghosting Psychology: Why Silence Speaks Volumes
Recruiter ghosting psychology reveals why candidates face silence after interviews. This post explores why companies ghost job applicants, ethical concerns, and how to interpret the silence.
Gap Analysis: Identify and Close Your Skill Gaps Effectively
Most candidates miss the exact skills that separate them from their target role. This overview shows how to map gaps, prioritize learning, and build a tangible plan to close them in months.