HyperAIHyperAI

Command Palette

Search for a command to run...

AI Companies Purchase Rare Books for Training Data

Global AI developers, spearheaded by Anthropic, have initiated a covert campaign to acquire rare and out-of-print physical books worldwide, scanning them for machine learning training before systematically destroying the originals. The operation, internally designated Project Panama, surged in early 2026 after AI firms recognized a critical bottleneck in generative AI development: internet data saturation. As synthetic content floods web repositories, leading to model collapse, pre-2022 printed texts have emerged as the final untapped source of pristine, human-verified linguistic data. To secure these archives, Anthropic and affiliated entities bypassed traditional academic licensing in favor of bulk commercial purchases. Intermediaries such as ISBNdb, Singapore-based 2077AI, and Canada’s Zoom Books facilitated the procurement, responding to sudden, high-volume orders from American, Dutch, German, and Spanish booksellers. The logistics involve industrial-scale binding removal and high-speed digitization, followed by pulp recycling. While the process extracts valuable text, it permanently eliminates physical copies, a practice that has alarmed librarians, authors, and cultural preservationists. Legally, the initiative operates within a contested gray area of United States copyright law. In June 2025, Federal Judge William Alsup ruled in Bartz v. Anthropic that scanning lawfully purchased books for AI training constitutes transformative fair use under Section 107, particularly when physical originals are destroyed and digital outputs remain proprietary. This precedent, later echoed in cases involving OpenAI and Meta, effectively insulates firms from infringement claims. Anthropic subsequently secured a 1.5 billion dollar settlement for earlier unlicensed scanning practices, a fraction of its projected 2026 annual revenue of 47 billion dollars. The data acquisition strategy has triggered significant industry friction. Critics, including White House technology advisor David Sacks and former executives like Ed Newton-Rex, condemn the model as intellectually extractive and culturally destructive. Conversely, prominent figures such as Elon Musk have advocated for non-destructive archival preservation. Meanwhile, the broader market remains skeptical of the strategy’s viability. Historically, humanity has published approximately 130 million unique books, yielding roughly 30 to 40 trillion tokens. Next-generation language models, however, are projected to require upwards of 100 trillion tokens, rendering mass book destruction an ultimately insufficient solution to AI data hunger. As legal frameworks struggle to adapt and global archives dwindle, the conflict between rapid technological scaling and heritage preservation is poised to intensify.

Related Links