TNW→ original

AI-компании скупают старые книги: последние данные для обучения без ИИ-слопа

AI-компании начали скупать старые бумажные книги как источник тренировочных данных: тексты, написанные до массового прихода генеративного ИИ, — последний крупный пул, не отравленный машинным «слопом». Дата-брокер ISBNdb предлагает лабораториям доступ к миллионам изданий, но чтобы их оцифровать, корешки книг срезают, а покупатели не раскрывают, кто оплатил заказ.

AI-processed from TNW; edited by Hamidun News
AI-компании скупают старые книги: последние данные для обучения без ИИ-слопа
Source: TNW. Collage: Hamidun News.
◐ Listen to article

AI labs have started buying up old paper books en masse to train models on them: texts published before the arrival of generative AI remain the last large pool of data uncontaminated by machine "slop." Access to the catalog is sold by data broker ISBNdb, while the labs themselves do not disclose who is funding the purchases — reported by The Next Web in July 2026.

Why are books without AI slop valuable?

Old books are valuable because they are guaranteed to be written by humans: any text published before the launch of ChatGPT in November 2022 physically cannot contain content generated by a neural network. It is precisely this "purity" that turns library shelves into a scarce resource for training models.

The problem lies in the extraction method. To digitize a book with an industrial scanner, its spine is cut off and the pages are run through an automatic feeder — in other words, it is physically destroyed. According to The Next Web, millions of volumes are involved, and the name of the lab placing the order does not appear in the deals.

  • ISBNdb — a data broker selling labs access to a catalog of books
  • November 2022 — the launch of ChatGPT, the dividing line between "clean" and contaminated data
  • Digitization is destructive: spines are cut off, paper books are destroyed
  • Millions of editions are in circulation
  • Buyers do not disclose which AI company paid for the order

Why is data running out?

Large labs have already effectively exhausted the open internet as a source of quality texts. Epoch AI analysts estimated back in 2024 that the stock of high-quality publicly available human data could run out in the second half of the 2020s, and the race for the remaining "clean" corpora has only intensified since then.

Books are a premium asset in this race. A book provides long, coherent, edited text with rich language, rather than fragmentary posts and comments, which is why a single volume is worth the quality that labs are willing to pay brokers like ISBNdb for.

What is data contamination?

Data contamination is a situation where training texts from the internet increasingly consist of the output of neural networks themselves rather than material written by humans. A model trained on its own "slop" gradually degrades: it loses diversity and accuracy. Researchers described this effect in the journal Nature back in 2024 and called it model collapse.

Hence the hunt for physical books. The internet has been rapidly filling up with AI text since 2022, and separating human from machine content on the web is becoming ever harder, whereas a book's publication date is a reliable marker that the text in front of you is pre-digital and human.

"The world's best AI training data is sitting on a shelf," is how, according to

The Next Web, data broker ISBNdb's pitch goes.

The legal side remains a gray area: buying and scanning other people's books to train commercial models is the subject of rights holders' lawsuits, and the anonymity of the buyers only heightens questions about transparency.

What this means

For the industry, this is a signal of structural scarcity: quality human data is finite, and everything created after 2022 is valued less and less because of the admixture of AI content. As models demand ever more text, the "pre-digital" heritage — old books, archives, newspapers — is turning into a strategic and physically finite resource.

ZK
Hamidun News
AI news without noise. Daily editorial selection from 50+ sources. A product by Zhemal Khamidun, Head of AI at Alpina Digital.

Want to stop reading about AI and start using it?

AI News is a curated feed of AI/tech news. Hamidun Academy teaches you to use AI systematically in your work.

What do you think?
Loading comments…