
DiscoBench shows search agents fail by searching, not asking
A new benchmark called DiscoBench reveals a critical failure mode in AI search agents: they perform worse when they attempt repeated searches on ambiguous queries instead of requesting clarification from users. The best-performing models achieve only 43 percent accuracy overall, while those that search iteratively without asking follow-up questions drop to 51.9 percent. Removing query ambiguity improves accuracy by up to 40 points, suggesting that agent design must prioritize interactive disambiguation over autonomous search loops. This finding reshapes expectations around agentic AI systems, indicating that production search agents need explicit uncertainty handling and user interaction protocols to function reliably.73























