WhiSPA: Semantically and Psychologically Aligned Whisper with Self-Supervised Contrastive and Student-Teacher Learning

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Rao, Rajath, Ganesan, Adithya, Kjell, Oscar, Luby, Jonah, Raghavan, Akshay, Feltman, Scott, Ringwald, Whitney, Boyd, Ryan L., Luft, Benjamin, Ruggero, Camilo, Ryant, Neville, Kotov, Roman, Schwartz, H. Andrew
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866918040757600256
author Rao, Rajath
Ganesan, Adithya
Kjell, Oscar
Luby, Jonah
Raghavan, Akshay
Feltman, Scott
Ringwald, Whitney
Boyd, Ryan L.
Luft, Benjamin
Ruggero, Camilo
Ryant, Neville
Kotov, Roman
Schwartz, H. Andrew
author_facet Rao, Rajath
Ganesan, Adithya
Kjell, Oscar
Luby, Jonah
Raghavan, Akshay
Feltman, Scott
Ringwald, Whitney
Boyd, Ryan L.
Luft, Benjamin
Ruggero, Camilo
Ryant, Neville
Kotov, Roman
Schwartz, H. Andrew
contents Current speech encoding pipelines often rely on an additional text-based LM to get robust representations of human communication, even though SotA speech-to-text models often have a LM within. This work proposes an approach to improve the LM within an audio model such that the subsequent text-LM is unnecessary. We introduce WhiSPA (Whisper with Semantic and Psychological Alignment), which leverages a novel audio training objective: contrastive loss with a language model embedding as a teacher. Using over 500k speech segments from mental health audio interviews, we evaluate the utility of aligning Whisper's latent space with semantic representations from a text autoencoder (SBERT) and lexically derived embeddings of basic psychological dimensions: emotion and personality. Over self-supervised affective tasks and downstream psychological tasks, WhiSPA surpasses current speech encoders, achieving an average error reduction of 73.4% and 83.8%, respectively. WhiSPA demonstrates that it is not always necessary to run a subsequent text LM on speech-to-text output in order to get a rich psychological representation of human communication.
format Preprint
id arxiv_https___arxiv_org_abs_2501_16344
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle WhiSPA: Semantically and Psychologically Aligned Whisper with Self-Supervised Contrastive and Student-Teacher Learning
Rao, Rajath
Ganesan, Adithya
Kjell, Oscar
Luby, Jonah
Raghavan, Akshay
Feltman, Scott
Ringwald, Whitney
Boyd, Ryan L.
Luft, Benjamin
Ruggero, Camilo
Ryant, Neville
Kotov, Roman
Schwartz, H. Andrew
Audio and Speech Processing
Artificial Intelligence
Computation and Language
Sound
Current speech encoding pipelines often rely on an additional text-based LM to get robust representations of human communication, even though SotA speech-to-text models often have a LM within. This work proposes an approach to improve the LM within an audio model such that the subsequent text-LM is unnecessary. We introduce WhiSPA (Whisper with Semantic and Psychological Alignment), which leverages a novel audio training objective: contrastive loss with a language model embedding as a teacher. Using over 500k speech segments from mental health audio interviews, we evaluate the utility of aligning Whisper's latent space with semantic representations from a text autoencoder (SBERT) and lexically derived embeddings of basic psychological dimensions: emotion and personality. Over self-supervised affective tasks and downstream psychological tasks, WhiSPA surpasses current speech encoders, achieving an average error reduction of 73.4% and 83.8%, respectively. WhiSPA demonstrates that it is not always necessary to run a subsequent text LM on speech-to-text output in order to get a rich psychological representation of human communication.
title WhiSPA: Semantically and Psychologically Aligned Whisper with Self-Supervised Contrastive and Student-Teacher Learning
topic Audio and Speech Processing
Artificial Intelligence
Computation and Language
Sound
url https://arxiv.org/abs/2501.16344