Benchmark reveals LLM agents struggle with long-term mental health prediction
Researchers have built BALMS, the first benchmark for evaluating how well LLM-based agents can extract mental health insights from continuous wearable data over extended periods. Current agentic systems excel at short-term retrieval tasks but fail to reason across longitudinal signals to forecast wellbeing outcomes with explainable evidence. This work matters because it exposes a critical gap between lab-ready LLM agents and real-world clinical deployment, where temporal reasoning and grounded predictions are non-negotiable. The benchmark establishes evaluation standards that will shape how the AI community approaches healthcare applications requiring sustained behavioral monitoring.
Modelwire context
ExplainerBALMS doesn't just measure agent performance on wearables; it isolates a specific failure mode: agents can retrieve isolated datapoints but cannot reason across weeks or months of behavioral signals to make grounded predictions. This temporal reasoning gap has been largely invisible in prior benchmarks focused on single-turn tasks.
This connects directly to the data generation framework from late August (the ACE lens paper), which highlighted how agentic data pipelines currently conflate construction with validation across domains. BALMS surfaces a concrete consequence: without longitudinal interaction traces and verification logic built into benchmark design, agents trained on fragmented short-term data cannot transfer to clinical settings where temporal coherence is non-negotiable. The wearable sensor foundation (HALO, also from this week) provides the hardware abstraction layer; BALMS now demands the reasoning layer on top of it.
If teams retrain existing agents on BALMS's longitudinal traces and close the gap to 75%+ accuracy within six months, the benchmark has identified a solvable architectural problem. If the gap persists or shrinks only marginally, it signals that current transformer designs may need fundamental changes to handle extended temporal dependencies in healthcare, which would reshape how foundation models approach clinical deployment.
Coverage we drew on
This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.
MentionsBALMS · LLM agents · wearable devices
Modelwire Editorial
This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.
Modelwire summarizes, we don’t republish. arXiv cs.CL originally reported this story as “BALMS: Benchmarking Agentic LLMs for Longitudinal Mental Health Sensing”. The full content lives on arxiv.org. If you’re a publisher and want a different summarization policy for your work, see our takedown page.