GOAT-SLM: A Spoken Language Model with Paralinguistic and Speaker Characteristic Awareness

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Chen, Hongjie, Li, Zehan, Song, Yaodong, Deng, Wenming, Yao, Yitong, Zhang, Yuxin, Lv, Hang, Zhu, Xuechao, Kang, Jian, Lian, Jie, Li, Jie, Wang, Chao, Song, Shuangyong, Li, Yongxiang, He, Zhongjiang, Li, Xuelong
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908466295078912
author Chen, Hongjie
Li, Zehan
Song, Yaodong
Deng, Wenming
Yao, Yitong
Zhang, Yuxin
Lv, Hang
Zhu, Xuechao
Kang, Jian
Lian, Jie
Li, Jie
Wang, Chao
Song, Shuangyong
Li, Yongxiang
He, Zhongjiang
Li, Xuelong
author_facet Chen, Hongjie
Li, Zehan
Song, Yaodong
Deng, Wenming
Yao, Yitong
Zhang, Yuxin
Lv, Hang
Zhu, Xuechao
Kang, Jian
Lian, Jie
Li, Jie
Wang, Chao
Song, Shuangyong
Li, Yongxiang
He, Zhongjiang
Li, Xuelong
contents Recent advances in end-to-end spoken language models (SLMs) have significantly improved the ability of AI systems to engage in natural spoken interactions. However, most existing models treat speech merely as a vehicle for linguistic content, often overlooking the rich paralinguistic and speaker characteristic cues embedded in human speech, such as dialect, age, emotion, and non-speech vocalizations. In this work, we introduce GOAT-SLM, a novel spoken language model with paralinguistic and speaker characteristic awareness, designed to extend spoken language modeling beyond text semantics. GOAT-SLM adopts a dual-modality head architecture that decouples linguistic modeling from acoustic realization, enabling robust language understanding while supporting expressive and adaptive speech generation. To enhance model efficiency and versatility, we propose a modular, staged training strategy that progressively aligns linguistic, paralinguistic, and speaker characteristic information using large-scale speech-text corpora. Experimental results on TELEVAL, a multi-dimensional evaluation benchmark, demonstrate that GOAT-SLM achieves well-balanced performance across both semantic and non-semantic tasks, and outperforms existing open-source models in handling emotion, dialectal variation, and age-sensitive interactions. This work highlights the importance of modeling beyond linguistic content and advances the development of more natural, adaptive, and socially aware spoken language systems.
format Preprint
id arxiv_https___arxiv_org_abs_2507_18119
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle GOAT-SLM: A Spoken Language Model with Paralinguistic and Speaker Characteristic Awareness
Chen, Hongjie
Li, Zehan
Song, Yaodong
Deng, Wenming
Yao, Yitong
Zhang, Yuxin
Lv, Hang
Zhu, Xuechao
Kang, Jian
Lian, Jie
Li, Jie
Wang, Chao
Song, Shuangyong
Li, Yongxiang
He, Zhongjiang
Li, Xuelong
Computation and Language
Artificial Intelligence
Sound
Audio and Speech Processing
Recent advances in end-to-end spoken language models (SLMs) have significantly improved the ability of AI systems to engage in natural spoken interactions. However, most existing models treat speech merely as a vehicle for linguistic content, often overlooking the rich paralinguistic and speaker characteristic cues embedded in human speech, such as dialect, age, emotion, and non-speech vocalizations. In this work, we introduce GOAT-SLM, a novel spoken language model with paralinguistic and speaker characteristic awareness, designed to extend spoken language modeling beyond text semantics. GOAT-SLM adopts a dual-modality head architecture that decouples linguistic modeling from acoustic realization, enabling robust language understanding while supporting expressive and adaptive speech generation. To enhance model efficiency and versatility, we propose a modular, staged training strategy that progressively aligns linguistic, paralinguistic, and speaker characteristic information using large-scale speech-text corpora. Experimental results on TELEVAL, a multi-dimensional evaluation benchmark, demonstrate that GOAT-SLM achieves well-balanced performance across both semantic and non-semantic tasks, and outperforms existing open-source models in handling emotion, dialectal variation, and age-sensitive interactions. This work highlights the importance of modeling beyond linguistic content and advances the development of more natural, adaptive, and socially aware spoken language systems.
title GOAT-SLM: A Spoken Language Model with Paralinguistic and Speaker Characteristic Awareness
topic Computation and Language
Artificial Intelligence
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2507.18119