The Decoder→ original

FineBooks от Hugging Face: лучший OCR для исторических книг стоит менее $2 за тысячу страниц

Hugging Face и EleutherAI запустили FineBooks — бенчмарк OCR для исторических книг. Из 14 протестированных открытых моделей лучшей оказалась dots.mocr: 97,6% точности при стоимости менее $2 за тысячу страниц. Этого достаточно для датасетов обучения ИИ, но команда признаёт — до стандартов академических транскрипций пока не дотягивает.

AI-processed from The Decoder; edited by Hamidun News
FineBooks от Hugging Face: лучший OCR для исторических книг стоит менее $2 за тысячу страниц
Source: The Decoder. Collage: Hamidun News.
◐ Listen to article

Hugging Face, together with EleutherAI, has published the results of the FineBooks project: the team tested 14 open-source OCR models on more than 2,000 pages of historical books. The best solution found — the dots.mocr model — achieves 97.6% character-level accuracy at a cost of less than two dollars per thousand pages.

Why the re-digitization of historical books is needed

Language models are trained on massive text corpora, a significant portion of which consists of digitized books from public archives. The problem is that most of these books were scanned in the 1990s and 2000s using primitive OCR software. The result is thousands of systematic errors: confused letters, broken words, missing lines, and unreadable abbreviations.

When such "dirty" OCR text enters a training dataset, the model learns from distorted patterns. This degrades generation quality — especially in domains where historical texts make up a significant share: law, medicine, classical literature, and the history of science.

FineBooks is an attempt by Hugging Face and EleutherAI to systematically address this problem. The team deliberately chose historical book pages as the most challenging subject for OCR: non-standard old typefaces, yellowed paper, uneven ink, and typographic defects from centuries-old publications. Before FineBooks, there was no systematic comparison of open-source OCR solutions specifically on historical material.

Which model came out on top

The testing covered 14 open-source OCR systems on a sample of more than 2,000 pages of historical publications. The winner — the dots.mocr model — demonstrated 97.6% character-level accuracy. This means fewer than 3 erroneous characters per 100 — a threshold at which text becomes suitable for building training datasets for large language models.

Key figures from the FineBooks project:

  • 14 open-source OCR models compared using a unified methodology
  • More than 2,000 historical book pages in the test sample
  • Leader — dots.mocr: 97.6% character-level accuracy
  • Processing cost — less than $2 per thousand pages

The economic dimension is critically important. At a cost of under $2 per thousand pages, reprocessing a corpus of one million pages costs less than $2,000 — a realistic budget for non-commercial AI research organizations working on creating open datasets.

Where accuracy still falls short

The authors of FineBooks honestly acknowledge limitations. The 97.6% figure is sufficient for preparing AI training data, but insufficient for academic transcriptions. Scholarly transcriptions require near-perfect reproduction of the original: specialists in history and textual criticism work with texts where every comma carries meaning, and the acceptable error threshold is an order of magnitude lower than for AI datasets.

"The results are good enough for AI training data, but have not yet reached the level required for scholarly transcriptions," as stated in the

FineBooks project materials, jointly published by Hugging Face and EleutherAI.

This is an important distinction: the study honestly identifies the boundary of applicability for dots.mocr, without claiming to be a universal solution. dots.mocr meets the needs of the AI industry, but the academic community is waiting for higher accuracy.

What this means

FineBooks gives the industry a concrete, actionable benchmark: dots.mocr is a cost-effective tool for large-scale re-digitization of historical archives for the purpose of training LLMs. According to Hugging Face and EleutherAI, the problem of "dirty" OCR text in training corpora is real and solvable right now — without waiting for expensive commercial solutions. FineBooks opens a path for millions of pages of historical books to finally become a full-fledged part of the training corpora for open language models.

ZK
Hamidun News
AI news without noise. Daily editorial selection from 50+ sources. A product by Zhemal Khamidun, Head of AI at Alpina Digital.

Want to stop reading about AI and start using it?

AI News is a curated feed of AI/tech news. Hamidun Academy teaches you to use AI systematically in your work.

What do you think?
Loading comments…