Beyond Speech and More: Investigating the Emergent Ability of Speech Foundation Models for Classifying Physiological Time-Series Signals

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Phukan, Orchid Chetia, Behera, Swarup Ranjan, Girish, Akhtar, Mohd Mujtaba, Buduru, Arun Balaji, Sharma, Rajesh
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917805072318464
author Phukan, Orchid Chetia
Behera, Swarup Ranjan
Girish
Akhtar, Mohd Mujtaba
Buduru, Arun Balaji
Sharma, Rajesh
author_facet Phukan, Orchid Chetia
Behera, Swarup Ranjan
Girish
Akhtar, Mohd Mujtaba
Buduru, Arun Balaji
Sharma, Rajesh
contents Despite being trained exclusively on speech data, speech foundation models (SFMs) like Whisper have shown impressive performance in non-speech tasks such as audio classification. This is partly because speech shares some common traits with audio, enabling SFMs to transfer effectively. In this study, we push the boundaries by evaluating SFMs on a more challenging out-of-domain (OOD) task: classifying physiological time-series signals. We test two key hypotheses: first, that SFMs can generalize to physiological signals by capturing shared temporal patterns; second, that multilingual SFMs will outperform others due to their exposure to greater variability during pre-training, leading to more robust, generalized representations. Our experiments, conducted for stress recognition using ECG (Electrocardiogram), EMG (Electromyography), and EDA (Electrodermal Activity) signals, reveal that models trained on SFM-derived representations outperform those trained on raw physiological signals. Among all models, multilingual SFMs achieve the highest accuracy, supporting our hypothesis and demonstrating their OOD capabilities. This work positions SFMs as promising tools for new uncharted domains beyond speech.
format Preprint
id arxiv_https___arxiv_org_abs_2410_12645
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Beyond Speech and More: Investigating the Emergent Ability of Speech Foundation Models for Classifying Physiological Time-Series Signals
Phukan, Orchid Chetia
Behera, Swarup Ranjan
Girish
Akhtar, Mohd Mujtaba
Buduru, Arun Balaji
Sharma, Rajesh
Audio and Speech Processing
Signal Processing
68T45
I.2.7
Despite being trained exclusively on speech data, speech foundation models (SFMs) like Whisper have shown impressive performance in non-speech tasks such as audio classification. This is partly because speech shares some common traits with audio, enabling SFMs to transfer effectively. In this study, we push the boundaries by evaluating SFMs on a more challenging out-of-domain (OOD) task: classifying physiological time-series signals. We test two key hypotheses: first, that SFMs can generalize to physiological signals by capturing shared temporal patterns; second, that multilingual SFMs will outperform others due to their exposure to greater variability during pre-training, leading to more robust, generalized representations. Our experiments, conducted for stress recognition using ECG (Electrocardiogram), EMG (Electromyography), and EDA (Electrodermal Activity) signals, reveal that models trained on SFM-derived representations outperform those trained on raw physiological signals. Among all models, multilingual SFMs achieve the highest accuracy, supporting our hypothesis and demonstrating their OOD capabilities. This work positions SFMs as promising tools for new uncharted domains beyond speech.
title Beyond Speech and More: Investigating the Emergent Ability of Speech Foundation Models for Classifying Physiological Time-Series Signals
topic Audio and Speech Processing
Signal Processing
68T45
I.2.7
url https://arxiv.org/abs/2410.12645