TimeSRL: Generalizable Time-Series Behavioral Modeling via Semantic RL-Tuned LLMs -- A Case Study in Mental Health

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fan, Yuang, Xu, Lilin, Wu, Millie, Nie, Jingping, Chen, Qingyu, Yang, Yuzhe, Zhang, Zhuo, Liu, Xin, Nepal, Subigya, Jiang, Xiaofan, Xu, Xuhai "Orson"
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916031831736320
author Fan, Yuang
Xu, Lilin
Wu, Millie
Nie, Jingping
Chen, Qingyu
Yang, Yuzhe
Zhang, Zhuo
Liu, Xin
Nepal, Subigya
Jiang, Xiaofan
Xu, Xuhai "Orson"
author_facet Fan, Yuang
Xu, Lilin
Wu, Millie
Nie, Jingping
Chen, Qingyu
Yang, Yuzhe
Zhang, Zhuo
Liu, Xin
Nepal, Subigya
Jiang, Xiaofan
Xu, Xuhai "Orson"
contents Longitudinal passive sensing enables continuous health prediction, yet models often fail under cross-dataset distribution shifts. Traditional ML overfits cohort-specific artifacts, while Large Language Models (LLMs) struggle to reason reliably over long, heterogeneous time-series. We introduce TimeSRL, a two-stage LLM framework that routes predictions through an explicit semantic bottleneck. The model first abstracts raw signals into high-level natural language, then predicts behavioral outcomes from these abstractions alone. This forces the model to reason over semantic concepts that we argue generalize better than raw numbers. We optimize this process end-to-end using Group Relative Policy Optimization (GRPO) with Reinforcement Learning from Verifiable Rewards (RLVR), learning outcome-aligned abstractions without gold intermediate annotations. Instantiated on mental-health prediction, TimeSRL achieves state-of-the-art performance on a benchmark designed to stress-test cross-cohort generalization under a rigorous leave-one-dataset-out (LOSO) protocol, reducing mean absolute error (MAE) over strong non-LLM ML and LLM baselines by 3.1--10.1% and 9.5--44.1% for anxiety, and 3.2--9.6% and 27.4--57.6% for depression (all $p$s<0.05). TimeSRL significantly outperforms prior methods in cross-benchmark transfer across different sensing pipelines, rivaling its own within-domain performance without target-domain fine-tuning. These results demonstrate that semantic abstractions are reusable and point to a new direction for generalizable behavior modeling via RL-tuned LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2605_21295
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle TimeSRL: Generalizable Time-Series Behavioral Modeling via Semantic RL-Tuned LLMs -- A Case Study in Mental Health
Fan, Yuang
Xu, Lilin
Wu, Millie
Nie, Jingping
Chen, Qingyu
Yang, Yuzhe
Zhang, Zhuo
Liu, Xin
Nepal, Subigya
Jiang, Xiaofan
Xu, Xuhai "Orson"
Machine Learning
Artificial Intelligence
Human-Computer Interaction
Longitudinal passive sensing enables continuous health prediction, yet models often fail under cross-dataset distribution shifts. Traditional ML overfits cohort-specific artifacts, while Large Language Models (LLMs) struggle to reason reliably over long, heterogeneous time-series. We introduce TimeSRL, a two-stage LLM framework that routes predictions through an explicit semantic bottleneck. The model first abstracts raw signals into high-level natural language, then predicts behavioral outcomes from these abstractions alone. This forces the model to reason over semantic concepts that we argue generalize better than raw numbers. We optimize this process end-to-end using Group Relative Policy Optimization (GRPO) with Reinforcement Learning from Verifiable Rewards (RLVR), learning outcome-aligned abstractions without gold intermediate annotations. Instantiated on mental-health prediction, TimeSRL achieves state-of-the-art performance on a benchmark designed to stress-test cross-cohort generalization under a rigorous leave-one-dataset-out (LOSO) protocol, reducing mean absolute error (MAE) over strong non-LLM ML and LLM baselines by 3.1--10.1% and 9.5--44.1% for anxiety, and 3.2--9.6% and 27.4--57.6% for depression (all $p$s<0.05). TimeSRL significantly outperforms prior methods in cross-benchmark transfer across different sensing pipelines, rivaling its own within-domain performance without target-domain fine-tuning. These results demonstrate that semantic abstractions are reusable and point to a new direction for generalizable behavior modeling via RL-tuned LLMs.
title TimeSRL: Generalizable Time-Series Behavioral Modeling via Semantic RL-Tuned LLMs -- A Case Study in Mental Health
topic Machine Learning
Artificial Intelligence
Human-Computer Interaction
url https://arxiv.org/abs/2605.21295