Dynamic Stress Detection: A Study of Temporal Progression Modelling of Stress in Speech

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lall, Vishakha, Liu, Yisi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910053381963776
author Lall, Vishakha
Liu, Yisi
author_facet Lall, Vishakha
Liu, Yisi
contents Detecting psychological stress from speech is critical in high-pressure settings. While prior work has leveraged acoustic features for stress detection, most treat stress as a static label. In this work, we model stress as a temporally evolving phenomenon influenced by historical emotional state. We propose a dynamic labelling strategy that derives fine-grained stress annotations from emotional labels and introduce cross-attention-based sequential models, a Unidirectional LSTM and a Transformer Encoder, to capture temporal stress progression. Our approach achieves notable accuracy gains on MuSE (+5%) and StressID (+18%) over existing baselines, and generalises well to a custom real-world dataset. These results highlight the value of modelling stress as a dynamic construct in speech.
format Preprint
id arxiv_https___arxiv_org_abs_2510_08586
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Dynamic Stress Detection: A Study of Temporal Progression Modelling of Stress in Speech
Lall, Vishakha
Liu, Yisi
Audio and Speech Processing
Artificial Intelligence
Computation and Language
Sound
Detecting psychological stress from speech is critical in high-pressure settings. While prior work has leveraged acoustic features for stress detection, most treat stress as a static label. In this work, we model stress as a temporally evolving phenomenon influenced by historical emotional state. We propose a dynamic labelling strategy that derives fine-grained stress annotations from emotional labels and introduce cross-attention-based sequential models, a Unidirectional LSTM and a Transformer Encoder, to capture temporal stress progression. Our approach achieves notable accuracy gains on MuSE (+5%) and StressID (+18%) over existing baselines, and generalises well to a custom real-world dataset. These results highlight the value of modelling stress as a dynamic construct in speech.
title Dynamic Stress Detection: A Study of Temporal Progression Modelling of Stress in Speech
topic Audio and Speech Processing
Artificial Intelligence
Computation and Language
Sound
url https://arxiv.org/abs/2510.08586