OpenAI researcher predicts $100B shift toward specialized training data

The scaling paradigm that powered recent LLM advances is hitting a wall. Andrew Ho, departing OpenAI to launch a data-focused startup, argues that raw compute increases alone no longer drive broad capability gains. Instead, models are narrowing into specialized performers, excelling at code and mathematics while losing ground elsewhere. This shift signals a fundamental constraint: labs must now invest heavily in curated, task-specific training datasets rather than simply feeding larger models more generic text. Ho's $100 billion prediction reflects a coming reallocation of AI R&D spending away from hardware and toward data engineering and collection infrastructure.
Modelwire context
Analyst takeThe more pointed detail here is the capability asymmetry Ho describes: models are not plateauing uniformly but are actually regressing in some domains while improving in others, which means the scaling critique isn't just about diminishing returns but about active distortion of model behavior.
This is largely disconnected from recent activity in our archive, as we have no prior coverage to anchor it to. It does, however, belong to a broader conversation that has been building across the industry around post-pretraining bottlenecks. The argument that curated data beats raw scale has been circulating in research circles for roughly two years, but Ho's move to commercialize that thesis is the notable step. A senior researcher leaving OpenAI to build infrastructure around a specific technical conviction is a signal worth tracking on its own terms, separate from whether the $100 billion figure is grounded in anything more than a confident round number.
Watch whether any of the major data labeling or synthetic data companies (Scale AI, Appen, or similar) announce partnerships or acqui-hire activity tied to Ho's new venture within the next six months. That would confirm the thesis is attracting capital, not just attention.
This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.
MentionsOpenAI · Andrew Ho · Adam Hunt · Cambridge · The Decoder
Modelwire Editorial
This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.
Modelwire summarizes, we don’t republish. The Decoder originally reported this story as “Ex-OpenAI researcher bets $100 billion will flow into training data because scaling alone won't cut it”. The full content lives on the-decoder.com. If you’re a publisher and want a different summarization policy for your work, see our takedown page.