Hugging Face and EleutherAI benchmark OCR for training data at scale

Hugging Face and EleutherAI have benchmarked OCR systems on historical texts, identifying a practical threshold for training data quality. The dots.mocr model achieves 97.6% character accuracy at scale, sufficient for language model pretraining but below scholarly standards. This work addresses a concrete bottleneck in synthetic data pipelines: digitized books remain a major training corpus, yet degraded OCR introduces systematic noise that compounds during model training. The finding matters because it quantifies the accuracy floor needed for large-scale document ingestion, helping teams decide whether to clean legacy scans or source fresh material.
Modelwire context
Skeptical readThe real question is whether 97.6% is a genuine floor or a marketing-friendly number. The summary doesn't say whether Hugging Face and EleutherAI tested this threshold across different model sizes, domains, or downstream tasks, or whether it only holds for the specific pretraining setup they used.
This connects directly to the data quality problem surfaced in 'Notes on the third era of slop' (early August). That piece identified a fracture between quality-first and scale-first strategies in the AI ecosystem. FineBooks is betting that teams will choose quality over volume when given a quantified trade-off, but the OCR accuracy floor only matters if downstream users actually adopt it. The inference optimization work from Baseten (same period) showed that speed and cost now rival capability as differentiators, which means teams under margin pressure may ignore this benchmark and use cheaper, dirtier data anyway.
If major labs (Anthropic, Meta, or xAI) publicly adopt the 97.6% threshold in their next pretraining reports, or explicitly reject it as too loose, that confirms whether this benchmark has real teeth. Otherwise, watch whether FineBooks' customer list includes anyone beyond Hugging Face ecosystem projects within the next six months. If adoption stays niche, the threshold was descriptive of one setup, not prescriptive for the field.
Coverage we drew on
- Notes on the third era of slop · Platformer
This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.
MentionsHugging Face · EleutherAI · FineBooks · dots.mocr
Modelwire Editorial
This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.
Modelwire summarizes, we don’t republish. The Decoder originally reported this story as “Old OCR text cripples language model training, and FineBooks wants to fix that at scale”. The full content lives on the-decoder.com. If you’re a publisher and want a different summarization policy for your work, see our takedown page.