BALMS: Benchmarking Agentic LLMs for Mental Health Sensing
Abstract
Ubiquitous wearable and mobile devices can collect large-scale, longitudinal multivariate behavioral and physiological signals, opening new opportunities for personalized monitoring of mental wellbeing. LLM-based agentic systems, which augment language models with reasoning, planning, and external tools, offer a promising paradigm for interpreting such longitudinal records and providing personalized, interactive health support. However, most existing healthcare agents are designed for clinical text and have not been systematically evaluated on daily-life multivariate sensor data, leaving the capabilities and limitations of agentic paradigms on this structured longitudinal modality unclear. We present a benchmark of LLM-based agentic systems for longitudinal wearable mental wellbeing tasks, spanning multiple datasets, using both closed and open source backbones, and three representative agentic frameworks. Through zero-shot and reasoning evaluation and sensitivity analyses over temporal length, we jointly report performance, cost, and latency, and distill design insights for agentic systems on longitudinal mobile health.