
SWE-INTERACT: Reimagining SWE Benchmarks as User-Driven Long-Horizon Coding Sessions
SWE-Interact reframes software engineering benchmarks around realistic developer workflows rather than autonomous task completion. Instead of handing agents complete specifications upfront, the testbed simulates iterative collaboration: a user simulator begins with vague requirements, inspects intermediate work, and progressively refines constraints. This shift matters because it exposes whether coding agents can handle requirement discovery, adapt to feedback loops, and build incrementally on their own output. The benchmark design reflects how actual development teams operate, making it a more honest stress test for production-ready coding systems than existing autonomous-only evaluations.62



























