Beyond the mean: Sequence analysis methods for clustering ordinal EMA data

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Tianyi, Smith, Anna L., Silva-Jones, Jillian R., Mendes, Wendy Berry, Whitehurst, Lauren N.
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913063307837440
author Wang, Tianyi
Smith, Anna L.
Silva-Jones, Jillian R.
Mendes, Wendy Berry
Whitehurst, Lauren N.
author_facet Wang, Tianyi
Smith, Anna L.
Silva-Jones, Jillian R.
Mendes, Wendy Berry
Whitehurst, Lauren N.
contents Ecological momentary assessment (EMA) ratings are widely used in studies of behavioral and psychological phenomena to capture real-time data in subjects' real-world environments. Because the data are collected repeatedly over the study period, they provide rich longitudinal rating profiles for each individual. However, the number of observations per subject is often large, while both sample size and sampling intensity can vary substantially across individuals, which complicates the analysis. In some settings, simplified summaries of individual profiles, such as averages computed across the study period, are used for downstream analyses, including regression-style modeling. Although such summaries can be convenient, they may fail to fully capture dynamic temporal patterns present in the complete longitudinal profiles. To address this, we borrow measures from sequence analysis that capture individual-level patterns over time and then applied principal component analysis (PCA) followed by $K$-means clustering to identify unobserved latent groups of individuals with similar profiles. We test our approach using simulated data from a categorical functional regression model and compare its performance with two commonly used methods for detecting unobserved group structures: latent class analysis (LCA), and latent transition analysis (LTA). Using EMA stress observations from a large sample of U.S. adults (Newman et al., 2024, 2025), we identify distinct latent stress profile groups and show that they improve characterization of the impact on cognitive performance.
format Preprint
id arxiv_https___arxiv_org_abs_2604_23834
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Beyond the mean: Sequence analysis methods for clustering ordinal EMA data
Wang, Tianyi
Smith, Anna L.
Silva-Jones, Jillian R.
Mendes, Wendy Berry
Whitehurst, Lauren N.
Methodology
Applications
Ecological momentary assessment (EMA) ratings are widely used in studies of behavioral and psychological phenomena to capture real-time data in subjects' real-world environments. Because the data are collected repeatedly over the study period, they provide rich longitudinal rating profiles for each individual. However, the number of observations per subject is often large, while both sample size and sampling intensity can vary substantially across individuals, which complicates the analysis. In some settings, simplified summaries of individual profiles, such as averages computed across the study period, are used for downstream analyses, including regression-style modeling. Although such summaries can be convenient, they may fail to fully capture dynamic temporal patterns present in the complete longitudinal profiles. To address this, we borrow measures from sequence analysis that capture individual-level patterns over time and then applied principal component analysis (PCA) followed by $K$-means clustering to identify unobserved latent groups of individuals with similar profiles. We test our approach using simulated data from a categorical functional regression model and compare its performance with two commonly used methods for detecting unobserved group structures: latent class analysis (LCA), and latent transition analysis (LTA). Using EMA stress observations from a large sample of U.S. adults (Newman et al., 2024, 2025), we identify distinct latent stress profile groups and show that they improve characterization of the impact on cognitive performance.
title Beyond the mean: Sequence analysis methods for clustering ordinal EMA data
topic Methodology
Applications
url https://arxiv.org/abs/2604.23834