New benchmark tests whether LLM agents can sustain long-term autonomy
Researchers have identified a critical gap in how LLM agents are evaluated. Current benchmarks test isolated, short-duration tasks in static settings, but real-world personal assistance demands sustained autonomy over weeks with shifting constraints and unannounced environmental changes. VibeLifeBench addresses this by introducing 200 long-horizon tasks that measure whether agents can maintain coherent planning, act proactively without prompting, and adapt to drift. This work signals growing recognition that capability measurement must evolve beyond single-turn performance to capture the persistence and self-directed decision-making required for deployed assistants.62



















