In the rapidly evolving landscape of artificial intelligence, data has become the world’s most contested commodity. As tech conglomerates exhaust the readily available, high-quality text on the open internet, they have turned to an unexpected source: physical books. However, because digitizing books at an industrial scale often involves physically destroying them—chopping off their spines to run pages through high-speed sheet feeders—the optics of this practice are disastrous. To shield themselves from public backlash, artificial intelligence companies have increasingly relied on opaque networks of third-party buyers, database managers, and book liquidators. Recently, this shadow supply chain was thrust into the spotlight when ISBNdb, a prominent online book database, quietly scrubbed its website of services explicitly designed to source and destroy physical books for Large Language Model (LLM) training. Main Facts: The Industrial Digitization and Destruction of Literature At the heart of the controversy is a simple, physical reality: to train cutting-edge LLMs, AI developers require trillions of tokens of high-quality, human-written text. While digital e-books exist, they are often protected by digital rights management (DRM) software, restrictive licensing agreements, and aggressive anti-piracy monitoring. Physical books, purchased legally on the secondary market, present a loophole. Under the "first-sale doctrine," a buyer is legally permitted to sell, donate, or destroy a physical copy of a copyrighted work. However, scanning a bound book page-by-page on a flatbed scanner is slow, labor-intensive, and economically unviable at the scale required for AI training. To automate the process, industrial scanning operations utilize high-speed sheet-fed scanners. This process requires "de-binding"—using industrial paper cutters to slice off the spine of the book, reducing a bound volume into a stack of loose, double-sided pages. Once fed through the scanner, the physical paper is typically discarded or sent to recycling facilities. [Physical Book Sourced] ➔ [Spine Guillotined/De-bound] ➔ [High-Speed Sheet Scanning] ➔ [OCR Text Processing] ➔ [AI Training Dataset] │ ▼ [Paper Sent to Recycle] This practice of literal book burning—or book shredding—has forced AI developers to use intermediary "rebuyers" to obfuscate their connection to the destruction of physical literature. ISBNdb, a platform widely used by libraries, researchers, and booksellers to retrieve book metadata, was recently exposed as one of these intermediaries, offering bulk-sourcing services tailored specifically for machine learning developers. Chronology of the Controversy The exposure of this clandestine supply chain unfolded over several months, driven by investigative reporting, independent bookseller observations, and digital archiving. +-----------------------------------------------------------------------------+ | CHRONOLOGY | +-----------------------------------------------------------------------------+ | | | [Early 2026] | | 404 Media publishes initial investigation exposing AI developers using | | front companies to purchase and destroy millions of physical books. | | | | [Mid 2026] | | Independent booksellers globally (notably in Australia) report bizarre, | | unprecedented bulk orders for rare and out-of-print titles. | | | | [Late July 2026] | | ISBNdb launches a dedicated landing page and blog post ("Reframing the | | Destruction Narrative") pitching bulk physical book sourcing for LLMs. | | | | [August 2026] | | Following intense public backlash and media inquiries, ISBNdb scrubs all | | AI-related sourcing pages from its website, claiming it was a "market test."| | | +-----------------------------------------------------------------------------+ Phase 1: The Initial Discovery The tech journalism outlet 404 Media broke the original story, revealing that AI companies were anonymously purchasing millions of physical books through middleman services. The investigation showed that these middlemen bought books in bulk from thrift stores, library sales, and liquidators, scanned them, and subsequently destroyed the physical copies. Phase 2: Booksellers Sound the Alarm Following the initial reports, Guardian Australia published an investigation revealing that independent and secondhand booksellers were noticing highly unusual purchasing patterns. Rather than buying popular bestsellers, mysterious online accounts were purchasing massive quantities of obscure, academic, and out-of-print books. Booksellers described these purchases as "waves of orders that did not fit with usual customer patterns," raising concerns that rare cultural artifacts were being permanently removed from circulation to be fed into scanners. Phase 3: The Intermediary Exposed In late July 2026, researchers and journalists discovered that ISBNdb had leaned entirely into this market. The database company had launched a dedicated landing page pitching itself as a "streamlined partner for sourcing printed books in bulk, tailored to your LLM training needs, delivered at the scale AI demands." Alongside the service page, ISBNdb published a blog post defending the practice of physical book destruction. Phase 4: The Retraction and Erasure Faced with immediate public outrage from authors, librarians, and copyright advocates, ISBNdb quietly deleted the LLM sourcing landing page and the accompanying blog post. However, digital archivists preserved the pages via the Internet Archive’s Wayback Machine, leaving a permanent record of the company’s venture into the AI data-harvesting market. Supporting Data and the Mechanics of "Book Migration" The commercial drive to scan physical books is fueled by a looming crisis in the AI industry: "model collapse." As the internet becomes saturated with AI-generated text, models trained on newer web data risk degrading in quality because they are learning from their own synthetic outputs. Physical books printed before the mid-2020s represent a pristine repository of pure human language, free from "AI slop." The Philosophy of "Value Migration" Before deleting its blog post, titled "Reframing the Destruction Narrative," ISBNdb attempted to provide a philosophical justification for the physical destruction of books. The company wrote: "The book is not destroyed. Its value has migrated. The paper returns to the material cycle; the knowledge enters the intellectual one." This corporate defense frames the physical book as a mere "container" for data, arguing that pulping a physical text is a form of recycling that upgrades the knowledge into a more accessible, digital format. However, critics point out several flaws in this argument: Loss of Artifacts: For rare, local, or out-of-print books, the physical copy destroyed may be one of only a few remaining in the world. Loss of Context: Marginalia, binding techniques, and historical physical context are entirely lost during high-speed scanning. The Hallucination Problem: Once ingested into an LLM, the book’s precise text is not preserved for public reading; instead, it is converted into mathematical weights. When asked to retrieve information from the book, the AI may hallucinate, distorting the very "knowledge" that supposedly migrated. Official Responses and Corporate Damage Control Following the removal of the controversial pages, ISBNdb attempted to distance itself from the bulk-sourcing business. When questioned by 404 Media regarding the abrupt deletion of their LLM training services, a representative for ISBNdb claimed: "The page was a test of market interest; no such service was ever brought to life." This response has met with deep skepticism from industry analysts. The presence of a fully realized landing page, integration with their existing book database infrastructure, and a deeply researched, philosophically defensive blog post suggest that the project was far more than a mere hypothetical test. Instead, critics argue it was an active business venture hastily aborted once the public realized the company was facilitating the destruction of physical literature. Meanwhile, the major AI developers suspected of utilizing these services—such as OpenAI, Meta, Google, and Anthropic—have remained silent on the matter. By using layers of anonymous shell companies and database intermediaries like ISBNdb, these tech giants maintain plausible deniability, avoiding headlines that connect their brands to the physical destruction of books. Implications for Culture, Copyright, and the Future of Books The revelation that physical books are being bought and destroyed systematically for AI training has profound implications across multiple sectors. The Threat to Out-of-Print and Rare Volumes When AI buyers sweep through secondhand book markets, they do not target common paperbacks. Instead, they seek out obscure non-fiction, local histories, academic monographs, and out-of-print texts to diversify their training data. Because many of these books were printed in limited runs, their physical destruction represents a permanent loss to human heritage. Unlike institutional libraries, which preserve physical volumes for posterity, the AI supply chain views these texts as disposable fuel. Book Category Value to AI Developers Risk Level Preserved in LLM? Popular Bestsellers Low (Already highly digitized/pirated online) Low Yes (Direct text reproduction likely) Academic Textbooks High (Dense, high-quality factual data) Medium No (Synthesized into model weights) Out-of-Print / Local History Extremely High (Unique vocabulary, rare facts) Critical (Often single-digit physical copies remaining) No (At risk of permanent erasure if destroyed) Intellectual Property and the Fair Use Battle The scanning of physical books highlights a massive legal loophole in copyright law. AI companies are currently facing numerous lawsuits from authors and publishers over the use of pirated digital book datasets (such as "Books3"). By purchasing physical books, AI companies can argue they are operating under the first-sale doctrine. However, copyright law distinguishes between owning a physical copy and reproducing it. The act of digitizing a physical book to train a commercial AI model stretches the boundaries of "fair use." If courts rule that scanning physical books for AI training constitutes copyright infringement, the entire physical-to-digital supply chain could face catastrophic legal liabilities. The Ethical Paradox of the "Intellectual Cycle" The controversy raises a fundamental ethical question about how society values physical objects versus digital utilities. Framing the destruction of physical books as a transition into the "intellectual cycle" suggests a future where physical libraries are obsolete, replaced by proprietary, commercial databases controlled by a handful of tech monopolies. For authors, the practice is a double insult: not only are their works ingested into systems designed to automate their profession, but the physical copies of their books—representing their tangible legacy—are guillotined and discarded in the process. As the AI industry continues its insatiable search for clean data, the battle over the written word is no longer confined to the digital realm. It is being fought in the dusty aisles of secondhand bookstores, where the physical survivors of the pre-digital era are quietly being bought up, sliced apart, and digitized into silence. Post navigation The Redemption of Rogue Trader: How Owlcat Games Navigated a Rocky Launch to Forge a CRPG Masterpiece