Multi-Turn Evaluation of Deep Research Agents Under Process-Level Feedback

Researchers have moved beyond single-turn evaluation of deep research agents by introducing a framework that tests whether these systems can improve iteratively under feedback. The work distinguishes between self-reflection, where agents revise autonomously, and process-level feedback, where external guidance targets specific research-strategy gaps. A novel method called Research Gap Inference infers where agents' reasoning breaks down by analyzing rubric mismatches. Early results show agents struggle with consistent improvement, incorporating and abandoning criteria at similar rates. This matters because production research agents will face real-world feedback loops, making iterative refinement a critical capability gap that current benchmarks have overlooked.
Modelwire context
ExplainerThe finding that agents incorporate and abandon rubric criteria at roughly equal rates is the buried signal here. That symmetry suggests the problem isn't attention or effort but something closer to an inability to maintain a coherent internal model of what 'improvement' means across turns.
This connects directly to the iOSWorld benchmark coverage from the same day, which exposed a parallel gap: current agent evaluations test isolated task completion rather than sustained, adaptive performance in messy real-world conditions. Both papers are essentially arguing that the field has been grading on the wrong curve. Where iOSWorld stresses agents across interconnected app environments and user-specific inference, this work stresses agents across time under corrective pressure. Together they sketch a more demanding picture of what production-grade agent capability actually requires, one that existing single-turn benchmarks cannot capture.
Watch whether any of the major deep research agent developers (Perplexity, OpenAI's Deep Research, Google's equivalent) publish internal evaluations using process-level feedback protocols within the next six months. Adoption of this framework by a production team would signal the field accepts iterative refinement as a first-class capability metric rather than a research curiosity.
Coverage we drew on
This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.
MentionsDeep Research Agents · Research Gap Inference
Modelwire Editorial
This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.
Modelwire summarizes, we don’t republish. The full content lives on arxiv.org. If you’re a publisher and want a different summarization policy for your work, see our takedown page.