AI labs turn to printed books to escape AI-generated training data
AI training pipelines face a quality crisis as models trained on internet-scale data increasingly absorb AI-generated content, degrading output fidelity. ISBNdb's pivot to sourcing printed books reveals a structural shift in how labs source training corpora: older, human-authored texts now command premium value precisely because they predate the AI-slop era. This exposes both a near-term data scarcity problem for frontier labs and a longer-term dependency on finite, non-renewable training material. The economics of book acquisition as a competitive moat signals that raw data authenticity has become a differentiator in model development.69
























