Back to feed
News Story
SSignal88
DeepTech深科技
1 sources

AI Companies Buy Rare Books for Training Data, Then Destroy Them

AI companies are acquiring rare physical books to scan and destroy for training data. Anthropic's 'Project Panama' aims to process 500,000 to 2 million books in six months. A US court ruled scanning legally purchased books for AI training is fair use, sparking debate.

SynthePulse Insight · AI deep reading

AI Companies' 'Out-of-Print Book' Famine: The Legal Destruction of Human Cultural Heritage

Version 1 · 1 source

As AI companies begin bulk purchasing and destroying rare physical books to obtain training data, a controversy over law, ethics, and the survival of civilization emerges.

  • Starting in 2026, AI companies have been bulk purchasing rare physical books through intermediaries, scanning them, and then destroying them for large model training.
  • Anthropic's 'Project Panama' planned to scan 500,000 to 2 million books within six months, with full confidentiality.
  • A U.S. court ruled that scanning legally purchased physical books for AI training constitutes 'transformative fair use,' legal but ethically controversial.
  • All human books can only provide 30-40 trillion tokens, while next-generation models require 100 trillion tokens, creating a massive data gap.
  • Intermediary ISBNdb publicly offered confidential procurement services, briefly deleted the page after exposure, then restored it, citing the court ruling as endorsement.
Open section navigationData Famine: From the Internet to Physical Books

Data Famine: From the Internet to Physical Books

Large model training relies on high-quality text, but internet corpora are being contaminated by AI-generated content, leading to 'model collapse'—models trained repeatedly on their own output degrade in quality. Thus, physical books published before 2022 have become the last clean data source, written by humans and free of AI-generated content.

Rare books, due to their limited print runs, have become a competitive advantage for AI companies. Starting in April 2026, U.S. booksellers noticed a surge in orders: buyers only provided ISBN lists, asked no questions about condition, and did not haggle; booksellers in the Netherlands and Germany also received bulk orders from Singapore's 2077AI and Canada's Zoom Books. Investigations revealed these orders came from AI companies, and the books were scanned and then destroyed.

'Project Panama': Legal but Controversial Scanning Operation

Court documents unsealed in early 2026 revealed Anthropic's 'Project Panama,' which began in early 2024, aiming to 'destructively scan every book in the world.' The company purchased physical books from used book dealers, cut off the spines, scanned them at high speed, recycled the paper, and kept digital copies for internal use only. The project planned to process 500,000 to 2 million books within six months, costing tens of millions of dollars.

Anthropic had previously faced legal risks for using pirated books, reaching a $1.5 billion settlement in July 2026 covering over 480,000 works, but that amount was only 3% of its annual revenue of $47 billion.

Legal Paradox: Destruction Is Safer

In June 2025, the U.S. District Court for the Northern District of California ruled in Bartz v. Anthropic that scanning legally purchased physical books for AI training constitutes 'transformative fair use.' The logic: the books were legally purchased (first sale doctrine), copying to digital form is transformative, and digital copies were not redistributed. Destroying the originals actually strengthens the 'substitution' argument, making the company's position in court more favorable.

Subsequently, other federal judges made similar rulings in cases involving OpenAI and Meta. However, Anthropic's earlier use of pirated scans was found to be infringing.

Shadow Supply Chain and Insatiable Appetite

Intermediary ISBNdb publicly offers AI training data procurement services, promising confidentiality, with order sizes ranging from 1,000 to 1 million books. After exposure, it briefly deleted the relevant page but soon restored it, citing the court ruling as legal backing. Other intermediaries like 2077AI and Zoom Books are also involved, but the identities of major clients remain unknown.

The approximately 130 million unique books in human history can only provide 30-40 trillion tokens, while next-generation models require 100 trillion tokens. Even scanning all books cannot meet the demand.

Ethical Controversy and Industry Reactions

Elon Musk has instructed his AI team to preserve rare books and scan them non-destructively. Former AI executive Ed Newton-Rex called this evidence of the power imbalance between tech giants and creators. White House tech advisor David Sacks criticized Anthropic for hypocrisy: using the world's output for free while opposing competitors using its output for training.

Booksellers have mixed feelings: they benefit economically but dislike seeing rare books turned into pulp. Authors are outraged that their creative work is used without permission, but the authors of books purchased by Anthropic may have been dead for decades, making copyright enforcement difficult.

Credibility boundary

This article is based on reporting by DeepTech, which cites sources such as The Washington Post and 404 Media, and references court documents. Key facts such as court rulings, settlement amounts, and project scales are sourced, but some details (e.g., identities of intermediary clients) remain speculative.

Insight takeaway

AI companies are legally but irreversibly consuming human cultural heritage to obtain clean training data, and legal loopholes and enormous data demands make this behavior difficult to curb, raising profound ethical and civilizational survival issues.

Primary report

DeepTech深科技

Primary source