Test-time scaling hits exploitation wall on open-ended tasks
Researchers benchmarked five test-time scaling families across open-ended generation tasks in medicine, law, finance, and creative writing, revealing a critical insight: exploration gains plateau quickly, but exploitation remains severely constrained. Unlike mathematics and code where verification is trivial, real-world domains expose a fundamental bottleneck. The finding reshapes how practitioners should allocate inference budgets, suggesting that generating more candidates yields diminishing returns while better refinement of existing outputs remains the frontier. This reframes the test-time compute scaling narrative from a solved problem to one requiring architectural rethinking for domains where ground truth is ambiguous.62

















