
Routing gap partly explained by label noise, not router failure
A new decomposition framework challenges the widely cited performance gap between learned routing systems and oracle-level performance in multi-model inference. The work reveals that much of this gap stems from label noise inherent in single-draw evaluation rather than fundamental router limitations. By separating reproducible specialist advantage from stochastic selection artifacts, researchers show that test-time sampling strategies like best-of-K can close the gap without requiring better routers. This reframes the routing optimization problem for practitioners and suggests current router underperformance may be overstated, with implications for cost-efficiency calculations in production LLM systems.58
























