AI firms are quietly buying and destroying millions of printed books to train their models
AI companies are purchasing large quantities of older printed books to use as training data for their models. They source these books through services like ISBNdb, which coordinates bulk purchases from secondary markets and emphasizes the books' pre‑2022 publication to avoid AI‑generated content contamination. The practice involves scanning and often destroying the physical books, raising concerns about transparency and the impact on the book trade.
- ▪ISBNdb facilitates bulk acquisitions of 1,000 to 1 million printed books, focusing on titles published before 2022 to ensure data cleanliness.
- ▪Sellers on platforms such as Alibris and Biblio report a sharp rise in large, random bulk orders that appear linked to AI training efforts.
- ▪Legal documents from a lawsuit involving Anthropic reveal plans to buy and scan millions of books, with many copies destroyed during high‑speed destructive scanning.
- ▪The process is covered by nondisclosure agreements, making it difficult to identify the specific AI firms behind the purchases.
TechSpot files mainly under tech. We currently carry 235 of its stories.
Story provenance
Source · retrieval · rights · ranking — open for full record
inspect →
Story provenance
Attribution is not the same as permission. This drawer separates discovery metadata, excerpts, WeSearch-generated summaries, reuse status, and whether the publisher receives the visit. Nothing here claims a legal grant the publisher has not made.
Record
| Original publisher | TechSpot |
| Canonical URL | https://www.techspot.com/news/113277-ai-firms-quietly-buying-destroying-millions-printed-books.html |
| Publication time | Wed, 29 Jul 2026 08:33 -0500 |
| Retrieval time | 2026-07-29T13:38:03.325Z |
| Last seen | 2026-07-29T13:38:03.325Z |
| Headline source | Publisher (no WeSearch rewrite) |
| Excerpt source | publisher body |
| Excerpt method | First ~120 words (~800 chars) of extracted publisher body, fair-use limited. |
| Summary | WeSearch · cerebras-chat (WeSearch summarizer) |
| Summary source text | contentText |
| Citation coverage | Summary is a WeSearch-generated derivative; primary citation is the original publisher URL. |
| Cluster | NAlQfXMtggOb · 1 stories |
| Cluster logic | Grouped by semantic title/content similarity across sources within a rolling window. Same-publisher template collisions are excluded from coverage comparison. |
| Ranking reason | Story pages are not engagement-ranked. Hub feeds use recency, with optional source-diversified chronological ordering (cap consecutive stories per source). No personalized ranking. |
| Publisher visit | Yes — open original |
| Substitutes article? | No — link-out required for full text |
Rights status (four layers)
WeSearch handling by dimension
| Indexing | May the item be indexed (stored, ranked, made findable)? | Allowed |
| Snippet | May a short excerpt of the publisher's text be shown? | Allowed |
| AI summary | May WeSearch generate its own short summary of the article? | Limited |
| Retrieval / RAG | May the content be exposed for third-party retrieval-augmented generation? | Not asserted |
| Model training | May the content be used to train AI models? | Not asserted |
| Commercial reuse | May the content be reused commercially? | Not permitted |
Basis: Derived from the published RSS/Atom feed. Contact: [email protected]. Reviewed: 2026-07-24.
Opening excerpt (first ~120 words) tap to expand
AI Industry books training AI firms are quietly buying and destroying millions of printed books to train their models Demand for older print books is surging as AI firms look for human-written content free from AI-generated text By Skye Jacobs July 29, 2026, 8:33 Add TechSpot Serving tech enthusiasts for over 25 years. TechSpot means tech analysis and advice you can trust. What we know so far: A market built for libraries and booksellers is now feeding AI training, as companies seek large volumes of printed books that predate the generative AI boom. ISBNdb, which says it maintains the world's largest book database, now helps AI companies source physical books in bulk.
…
Excerpt limited to ~120 words for fair-use compliance. The full article is at TechSpot.