
Multi-Turn Evaluation of Deep Research Agents Under Process-Level Feedback
Researchers have moved beyond single-turn evaluation of deep research agents by introducing a framework that tests whether these systems can improve iteratively under feedback. The work distinguishes between self-reflection, where agents revise autonomously, and process-level feedback, where external guidance targets specific research-strategy gaps. A novel method called Research Gap Inference infers where agents' reasoning breaks down by analyzing rubric mismatches. Early results show agents struggle with consistent improvement, incorporating and abandoning criteria at similar rates. This matters because production research agents will face real-world feedback loops, making iterative refinement a critical capability gap that current benchmarks have overlooked.62



























