Paralinguistics-Enhanced Large Language Modeling of Spoken Dialogue

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lin, Guan-Ting, Shivakumar, Prashanth Gurunath, Gandhe, Ankur, Yang, Chao-Han Huck, Gu, Yile, Ghosh, Shalini, Stolcke, Andreas, Lee, Hung-yi, Bulyko, Ivan
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909075514589184
author Lin, Guan-Ting
Shivakumar, Prashanth Gurunath
Gandhe, Ankur
Yang, Chao-Han Huck
Gu, Yile
Ghosh, Shalini
Stolcke, Andreas
Lee, Hung-yi
Bulyko, Ivan
author_facet Lin, Guan-Ting
Shivakumar, Prashanth Gurunath
Gandhe, Ankur
Yang, Chao-Han Huck
Gu, Yile
Ghosh, Shalini
Stolcke, Andreas
Lee, Hung-yi
Bulyko, Ivan
contents Large Language Models (LLMs) have demonstrated superior abilities in tasks such as chatting, reasoning, and question-answering. However, standard LLMs may ignore crucial paralinguistic information, such as sentiment, emotion, and speaking style, which are essential for achieving natural, human-like spoken conversation, especially when such information is conveyed by acoustic cues. We therefore propose Paralinguistics-enhanced Generative Pretrained Transformer (ParalinGPT), an LLM that utilizes text and speech modalities to better model the linguistic content and paralinguistic attributes of spoken dialogue. The model takes the conversational context of text, speech embeddings, and paralinguistic attributes as input prompts within a serialized multitasking multimodal framework. Specifically, our framework serializes tasks in the order of current paralinguistic attribute prediction, response paralinguistic attribute prediction, and response text generation with autoregressive conditioning. We utilize the Switchboard-1 corpus, including its sentiment labels as the paralinguistic attribute, as our spoken dialogue dataset. Experimental results indicate the proposed serialized multitasking method outperforms typical sequence classification techniques on current and response sentiment classification. Furthermore, leveraging conversational context and speech embeddings significantly improves both response text generation and sentiment prediction. Our proposed framework achieves relative improvements of 6.7%, 12.0%, and 3.5% in current sentiment accuracy, response sentiment accuracy, and response text BLEU score, respectively.
format Preprint
id arxiv_https___arxiv_org_abs_2312_15316
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Paralinguistics-Enhanced Large Language Modeling of Spoken Dialogue
Lin, Guan-Ting
Shivakumar, Prashanth Gurunath
Gandhe, Ankur
Yang, Chao-Han Huck
Gu, Yile
Ghosh, Shalini
Stolcke, Andreas
Lee, Hung-yi
Bulyko, Ivan
Computation and Language
Audio and Speech Processing
Large Language Models (LLMs) have demonstrated superior abilities in tasks such as chatting, reasoning, and question-answering. However, standard LLMs may ignore crucial paralinguistic information, such as sentiment, emotion, and speaking style, which are essential for achieving natural, human-like spoken conversation, especially when such information is conveyed by acoustic cues. We therefore propose Paralinguistics-enhanced Generative Pretrained Transformer (ParalinGPT), an LLM that utilizes text and speech modalities to better model the linguistic content and paralinguistic attributes of spoken dialogue. The model takes the conversational context of text, speech embeddings, and paralinguistic attributes as input prompts within a serialized multitasking multimodal framework. Specifically, our framework serializes tasks in the order of current paralinguistic attribute prediction, response paralinguistic attribute prediction, and response text generation with autoregressive conditioning. We utilize the Switchboard-1 corpus, including its sentiment labels as the paralinguistic attribute, as our spoken dialogue dataset. Experimental results indicate the proposed serialized multitasking method outperforms typical sequence classification techniques on current and response sentiment classification. Furthermore, leveraging conversational context and speech embeddings significantly improves both response text generation and sentiment prediction. Our proposed framework achieves relative improvements of 6.7%, 12.0%, and 3.5% in current sentiment accuracy, response sentiment accuracy, and response text BLEU score, respectively.
title Paralinguistics-Enhanced Large Language Modeling of Spoken Dialogue
topic Computation and Language
Audio and Speech Processing
url https://arxiv.org/abs/2312.15316