Synthetic Audio Helps for Cognitive State Tasks

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Soubki, Adil, Murzaku, John, Zeng, Peter, Rambow, Owen
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866929709436108800
author Soubki, Adil
Murzaku, John
Zeng, Peter
Rambow, Owen
author_facet Soubki, Adil
Murzaku, John
Zeng, Peter
Rambow, Owen
contents The NLP community has broadly focused on text-only approaches of cognitive state tasks, but audio can provide vital missing cues through prosody. We posit that text-to-speech models learn to track aspects of cognitive state in order to produce naturalistic audio, and that the signal audio models implicitly identify is orthogonal to the information that language models exploit. We present Synthetic Audio Data fine-tuning (SAD), a framework where we show that 7 tasks related to cognitive state modeling benefit from multimodal training on both text and zero-shot synthetic audio data from an off-the-shelf TTS system. We show an improvement over the text-only modality when adding synthetic audio data to text-only corpora. Furthermore, on tasks and corpora that do contain gold audio, we show our SAD framework achieves competitive performance with text and synthetic audio compared to text and gold audio.
format Preprint
id arxiv_https___arxiv_org_abs_2502_06922
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Synthetic Audio Helps for Cognitive State Tasks
Soubki, Adil
Murzaku, John
Zeng, Peter
Rambow, Owen
Sound
Artificial Intelligence
Computation and Language
Machine Learning
The NLP community has broadly focused on text-only approaches of cognitive state tasks, but audio can provide vital missing cues through prosody. We posit that text-to-speech models learn to track aspects of cognitive state in order to produce naturalistic audio, and that the signal audio models implicitly identify is orthogonal to the information that language models exploit. We present Synthetic Audio Data fine-tuning (SAD), a framework where we show that 7 tasks related to cognitive state modeling benefit from multimodal training on both text and zero-shot synthetic audio data from an off-the-shelf TTS system. We show an improvement over the text-only modality when adding synthetic audio data to text-only corpora. Furthermore, on tasks and corpora that do contain gold audio, we show our SAD framework achieves competitive performance with text and synthetic audio compared to text and gold audio.
title Synthetic Audio Helps for Cognitive State Tasks
topic Sound
Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2502.06922