AI guidance shifts from chat to autonomous agents in one year

Ethan Mollick's evolving guidance on AI tool selection reveals a fundamental market shift from chat interfaces toward agentic systems capable of autonomous multi-hour workflows. Within a year, the landscape has moved decisively beyond conversational models like ChatGPT and Claude toward agents that execute complex tasks end-to-end. Google's absence from current recommendations signals competitive pressure, while the prominence of o3, Claude 4 Opus, and Gemini 2.5 Pro reflects which vendors have successfully built reasoning and autonomy capabilities. This trajectory matters for practitioners: the value proposition has moved from augmentation to delegation, reshaping how teams should evaluate and deploy AI infrastructure.
Modelwire context
Analyst takeThe more pointed observation in Mollick's framing is that this isn't just a capability ranking but a workflow architecture recommendation: he's telling practitioners to stop optimizing for the best chat response and start evaluating which systems can hold state and execute autonomously over extended sessions. That's a different procurement question than most teams are currently asking.
This is largely disconnected from recent activity in our archive, as we have no prior coverage to anchor it to. It belongs to a broader conversation about the organizational shift from AI as a productivity add-on to AI as a delegated worker, a transition that has been building across enterprise software and developer tooling circles throughout 2025 and into 2026. Mollick's practitioner-facing framing is notable precisely because it bypasses vendor positioning and reflects actual workflow adoption patterns.
Watch whether Google responds to its absence from Mollick's recommendations with a concrete Gemini agent product update before Q4 2026. If Gemini 2.5 Pro reappears in his next revision alongside o3 and Claude 4 Opus for agentic tasks specifically, that would confirm Google has closed the autonomy gap rather than just the reasoning benchmark gap.
This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.
MentionsEthan Mollick · ChatGPT · Claude · Gemini · o3 · Claude 4 Opus
Modelwire Editorial
This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.
Modelwire summarizes, we don’t republish. Simon Willison originally reported this story as “An opinionated guide to which AI to use to do stuff”. The full content lives on simonwillison.net. If you’re a publisher and want a different summarization policy for your work, see our takedown page.