Amazon, once an online bookseller, is destroying rare books to train AI models
Rare books are incredibly valuable for training LLMs, since these models have already trained on whatever's available online.
Kiran Ch
Contributor
There is a poetic, almost grotesque irony in the fact that Amazon—a company that started its trillion-dollar ascent out of a Bellevue garage packing paperbacks into cardboard mailers—is now taking industrial guillotine blades to rare, out-of-print physical books. The objective isn't salvage or archival preservation; it is the raw, unsentimental extraction of high-entropy tokens. Machine learning labs have hit the public web's scraping limit, and in their panic over model collapse, they are sending vintage paper through high-speed document feeders to feed the next generation of multimodal transformers.
Let’s be entirely clear about what is happening behind closed doors at Big Tech AI divisions. The "data wall" isn't a theoretical bottleneck scheduled for 2026; it is here right now. The public web is polluted with SEO sludge, regurgitated AI outputs, and litigious paywalls. To prevent models from choking on their own synthetic exhaust, frontier AI developers need pristine, untouched human reasoning. And the densest, most uncorrupted repository of human thought happens to sit inside physical books printed before the internet existed. To digitize them fast enough to hit compute deadlines, you don't use white-glove archival flatbeds. You slice off the spine, toss the binding in the trash, and run the loose leaves through an optical scanner at 150 pages per minute.
Join 15,000+ tech leaders
Get instant alerts on the most critical AI breakthroughs on our WhatsApp channel. No spam, just pure alpha.
Key Takeaways
- The Pretraining Plateau: Frontier AI labs have exhausted crawlable high-quality web text, forcing them into physical-world data scavenging to secure novel, human-authored tokens.
- Destructive Scanning at Scale: Non-destructive planetary scanning is far too slow for training dataset deadlines, leading to the industrial decapitation of rare and out-of-print books for automated optical character recognition (OCR).
- Synthetic Data Skepticism: The aggressive pivot to offline historical print proves that self-consuming synthetic loops cannot solve the loss-curve stagnation currently hitting large foundation models.
- The First-Sale Legal Loophole: Big Tech is exploiting physical ownership frameworks to ingest out-of-print and orphan works, effectively privatizing public cultural heritage into closed-source model weights.
The Token Drought and the Scent of Sliced Paper
If you track the scaling laws laid down by Chinchilla and subsequent compute-optimal training research, you know that parameter growth demands a proportional expansion of unique, high-quality tokens. For the last five years, AI labs treated the open internet like an infinite buffet. They hoovered Common Crawl, scraped Reddit's entire back catalog, indexed every public GitHub repository, and pirated shadow libraries via scraped torrent trackers. But that well has run dry. The returns on scraping modern web pages are yielding negative marginal value because the internet is increasingly populated by lower-tier LLM hallucinations.
When you feed an LLM its own synthetic output over successive training generations, the statistical tails of the probability distribution decay. You get variance collapse. The model forgets how to reason across edge cases, its vocabulary shrinks, and its output homogenizes into corporate sludge. To push past the current frontier benchmark plateaus, labs need long-tail human reasoning: specialized 19th-century mechanical treatises, obscure mid-century academic journals, long-forgotten regional historical texts, and deep philosophical treatises that never received an ISBN, let alone a Kindle release.
These texts hold pristine lexical variance, complex grammatical structures, and factual dense chains of logic that exist nowhere in the digital footprint. For a company like Amazon, which commands an unmatched logistics network capable of acquiring physical inventory at scale, physical print is the ultimate untapped moat. They aren't just buying books to sell them to readers; they are buying them as consumable raw fuel for GPU clusters.
The Industrial Meat Grinder: Destructive Scanning Pipelines
Archivists and librarians approach rare books with humidity-controlled vaults, overhead planetary cameras, and gentle v-shaped cradles that photograph pages without stressing centuries-old glue. It takes hours to scan a single volume carefully. If your goal is to train a multi-trillion-parameter model before your next quarterly earnings call, that meticulous cadence is an engineering failure. You need millions of tokens per hour, not per week.
The solution adopted by large-scale ingestion operations is destructive digitization. An industrial blade shears off the spine of a vintage book in a fraction of a second. The detached pages are loaded into high-capacity automated document feeders (ADFs), which rip the sheets through dual-sensor optical scanners with zero regard for paper tears, ink degradation, or historical value. Once the raw TIF or PNG files hit the S3 bucket, the shredded paper is dropped straight into the recycling bin.
From an algorithmic perspective, what comes next is a brutal data pipeline. These raw scans run through vision-language processing units tasked with deskewing, removing bleed-through artifacts from historical inks, handling obsolete typography like the long 's', and extracting semantic layout structure. The pipeline converts physical artifacts into clean, tokenized JSONL files ready for pretraining. When critics argue that tech giants are disconnected from the societal downstream of their tooling, they often miss this physical layer of destruction. We have reached a point where the Anthropic CEO says AI backlash is ‘fundamentally a crisis of trust’, and watching tech empires physically pulp historical literature to train predictive text engines validates every bit of that public cynicism.
"We have mined the live internet to exhaustion, processed every accessible open-source repository, and now we are feeding centuries of irreplaceable physical print into industrial paper-cutters so an AI cluster can output slightly better corporate summaries."
The Dark Irony of the Everything Store’s Origins
You cannot ignore the poetic tragedy at the heart of this operational pivot. Amazon built its empire on the premise of democratizing access to books. In 1995, its competitive advantage was offering the "Earth’s Biggest Bookstore"—a digital catalog that allowed a reader in rural Idaho to order an obscure work of historical non-fiction that no local brick-and-mortar shop could afford to stock. Books were the wedge that gave Amazon the cash flow and customer base to build AWS, build the global freight infrastructure, and eventually construct world-class AI supercomputers.
Thirty years later, the books are no longer being preserved or distributed to human minds. They are being acquired systematically through secondary marketplaces, estate sales, and clearance lots to be rendered down into statistical matrices. It reflects a total shift in how technology platforms view human culture: not as ideas meant to be read, debated, and preserved across generations, but as unstructured training tokens awaiting ingestion.
This extractive logic explains a great deal about the broader friction between Silicon Valley and the public. Consumers are pushing back against heavy-handed algorithmic products across the board, in part because the value exchange has grown so lopsided. It’s the very reason why people arent buying Mark Zuckerberg’s AI future: the industry asks users and culture to surrender everything—their data, their creative output, and now their historical physical artifacts—in exchange for black-box software that attempts to displace the very humans who created the training data.
Synthetic Data Failure and the Realities of Chinchilla Scaling
For the past eighteen months, AI research papers have boasted that synthetic data—generating training tokens using stronger models like Claude 3.5 Sonnet or GPT-4o—would permanently eliminate the human data bottleneck. "We will just have AI teach AI," the venture capital pitches claimed. In practice, production-grade model training has exposed massive structural issues with purely synthetic data pipelines.
Synthetic data is fantastic for narrow instruction-tuning, mathematical step-by-step reinforcement, and code optimization tasks. But for foundational pretraining—where the base model acquires its broad worldview, cross-disciplinary reasoning faculties, and linguistic intuition—synthetic data acts like distilled water. It lacks the complex mineral content, historical context, idiosyncrasies, and organic messiness of true human discourse. When you train purely on synthetic datasets, models develop blind spots, parrot existing biases with higher confidence, and lose the ability to generalize across novel real-world domains.
That is why physical books are the holy grail of pretraining. An out-of-print 1940s manual on naval engineering or an 1880s geographical survey of South America contains dense, highly specialized vocabularies and logical progressions that no contemporary digital author is writing today. By destroying and digitizing these books, labs are effectively injecting high-purity, unpolluted human cognition directly into the base weights, giving them an algorithmic edge that pure compute cannot replicate.
Legal Loopholes, First Sale Doctrine, and Digital Alexandria
How does a corporation legally justify the systematic acquisition and destruction of copyrighted or out-of-print books for AI training? They lean hard on the First Sale Doctrine and the muddy legal waters of digital transformation. Under United States copyright law, once you buy a physical copy of a book, you own that specific physical object. You have the undisputed legal right to sell it, lend it, throw it in the fire, or slice its spine with an industrial blade.
The copyright infringement question only enters the conversation when those scanned pages are converted into digital files and used to train commercial models. AI developers argue that training is protected under Fair Use—that the model is merely "learning" statistical associations across billions of data points rather than distributing unauthorized digital copies. Because the physical book is destroyed during the process, there is no lingering unauthorized physical duplicate circulating in the market, complicating traditional property claims.
This creates a deeply troubling dynamic for out-of-print and orphan works—books where the copyright holder is dead or untraceable, yet the material has not formally entered the public domain. Instead of these works eventually being digitized by public archives like the Internet Archive or local libraries for universal public access, they are being acquired by private capital, digitized behind closed enterprise doors, and locked inside proprietary model weights. Much like the industry's messy experiments with tracking output origins—where engineers debate how Claude’s new watermarks will work to prevent copyright contamination—the upstream intake of physical history remains completely untracked, untraceable, and permanently privatized.
We are watching the quiet privatization of historical epistemology. The physical copies disappear into scrap heaps; the digital tokens vanish into proprietary multi-layered neural networks. The world gains an API endpoint that can answer questions with impressive speed, but loses the tangible, open, and verifiable historical record that made that knowledge possible in the first place.
Frequently Asked Questions
Why are AI companies targeting physical books instead of using internet data?
AI developers have largely exhausted the crawlable, high-quality text available on the open web. Furthermore, the modern internet is saturated with low-quality, SEO-driven content and synthetic AI-generated text, which degrades model performance. Physical books, particularly older or out-of-print titles, contain pristine, dense, and uncorrupted human reasoning that offers high-value tokens impossible to find online.
Why do companies destroy the books instead of scanning them carefully?
Non-destructive scanning methods—such as using overhead planetary cameras and manual page
Supercharge Your Workflow with Claude AI
The AI assistant used by 100K+ professionals. Write, code, analyse — all in one place.