Over 27 percent of web pretraining data is now AI-generated text
A large-scale empirical study reveals that over 27% of web text used to train language models is now AI-generated, with that share climbing toward 31% by mid-2026. Researchers pretrained 800 models to measure how this unlabeled synthetic data affects learning curves and generalization. The key finding: AI tokens initially improve loss on human text for data-constrained models, but gains plateau quickly, suggesting diminishing returns and potential long-term risks for model quality as the web becomes increasingly synthetic. This work quantifies a critical infrastructure challenge facing the pretraining pipeline.
Modelwire context
Analyst takeThe study doesn't just confirm AI-generated text is flooding training corpora; it measures the exact point where adding more synthetic data stops helping and starts hurting. That plateau matters because it reframes the problem from 'how much synthetic data can we use' to 'we've already hit the wall on quantity, so what now'.
This directly backstops OpenAI's September admission that AI-generated content is degrading internet quality. The feedback loop they described (training on synthetic text, releasing models that generate more synthetic text) now has empirical cost attached: diminishing returns on model performance. That creates pressure on the licensing model Google is piloting with publishers. If synthetic data stops improving models around the 31 percent threshold, frontier labs lose their justification for training on unlabeled web text and must shift to licensed or proprietary corpora. The Wuhan court's decision to price token consumption into copyright damages becomes more relevant in this context, because the economics of licensing versus scraping flip when free synthetic data stops yielding gains.
Monitor whether Anthropic, Meta, or other labs announce shifts toward licensed training data or proprietary synthetic corpora in Q4 2026. If the major players begin licensing content at scale rather than relying on web scraping, that confirms this paper's findings are driving real procurement decisions. Conversely, if they continue training on raw web text despite the plateau, that signals the quality loss is acceptable to them relative to licensing costs.
Coverage we drew on
This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.
MentionsPangram · FineWeb
Modelwire Editorial
This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.
Modelwire summarizes, we don’t republish. arXiv cs.CL originally reported this story as “How Much Is an AI Token Worth? Scaling Laws for Wild AI-Generated Web Text”. The full content lives on arxiv.org. If you’re a publisher and want a different summarization policy for your work, see our takedown page.