Paralinguistics-Enhanced Large Language Modeling of Spoken Dialogue
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909075514589184 |
|---|---|
| author | Lin, Guan-Ting Shivakumar, Prashanth Gurunath Gandhe, Ankur Yang, Chao-Han Huck Gu, Yile Ghosh, Shalini Stolcke, Andreas Lee, Hung-yi Bulyko, Ivan |
| author_facet | Lin, Guan-Ting Shivakumar, Prashanth Gurunath Gandhe, Ankur Yang, Chao-Han Huck Gu, Yile Ghosh, Shalini Stolcke, Andreas Lee, Hung-yi Bulyko, Ivan |
| contents | Large Language Models (LLMs) have demonstrated superior abilities in tasks such as chatting, reasoning, and question-answering. However, standard LLMs may ignore crucial paralinguistic information, such as sentiment, emotion, and speaking style, which are essential for achieving natural, human-like spoken conversation, especially when such information is conveyed by acoustic cues. We therefore propose Paralinguistics-enhanced Generative Pretrained Transformer (ParalinGPT), an LLM that utilizes text and speech modalities to better model the linguistic content and paralinguistic attributes of spoken dialogue. The model takes the conversational context of text, speech embeddings, and paralinguistic attributes as input prompts within a serialized multitasking multimodal framework. Specifically, our framework serializes tasks in the order of current paralinguistic attribute prediction, response paralinguistic attribute prediction, and response text generation with autoregressive conditioning. We utilize the Switchboard-1 corpus, including its sentiment labels as the paralinguistic attribute, as our spoken dialogue dataset. Experimental results indicate the proposed serialized multitasking method outperforms typical sequence classification techniques on current and response sentiment classification. Furthermore, leveraging conversational context and speech embeddings significantly improves both response text generation and sentiment prediction. Our proposed framework achieves relative improvements of 6.7%, 12.0%, and 3.5% in current sentiment accuracy, response sentiment accuracy, and response text BLEU score, respectively. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2312_15316 |
| institution | arXiv |
| publishDate | 2023 |
| record_format | arxiv |
| spellingShingle | Paralinguistics-Enhanced Large Language Modeling of Spoken Dialogue Lin, Guan-Ting Shivakumar, Prashanth Gurunath Gandhe, Ankur Yang, Chao-Han Huck Gu, Yile Ghosh, Shalini Stolcke, Andreas Lee, Hung-yi Bulyko, Ivan Computation and Language Audio and Speech Processing Large Language Models (LLMs) have demonstrated superior abilities in tasks such as chatting, reasoning, and question-answering. However, standard LLMs may ignore crucial paralinguistic information, such as sentiment, emotion, and speaking style, which are essential for achieving natural, human-like spoken conversation, especially when such information is conveyed by acoustic cues. We therefore propose Paralinguistics-enhanced Generative Pretrained Transformer (ParalinGPT), an LLM that utilizes text and speech modalities to better model the linguistic content and paralinguistic attributes of spoken dialogue. The model takes the conversational context of text, speech embeddings, and paralinguistic attributes as input prompts within a serialized multitasking multimodal framework. Specifically, our framework serializes tasks in the order of current paralinguistic attribute prediction, response paralinguistic attribute prediction, and response text generation with autoregressive conditioning. We utilize the Switchboard-1 corpus, including its sentiment labels as the paralinguistic attribute, as our spoken dialogue dataset. Experimental results indicate the proposed serialized multitasking method outperforms typical sequence classification techniques on current and response sentiment classification. Furthermore, leveraging conversational context and speech embeddings significantly improves both response text generation and sentiment prediction. Our proposed framework achieves relative improvements of 6.7%, 12.0%, and 3.5% in current sentiment accuracy, response sentiment accuracy, and response text BLEU score, respectively. |
| title | Paralinguistics-Enhanced Large Language Modeling of Spoken Dialogue |
| topic | Computation and Language Audio and Speech Processing |
| url | https://arxiv.org/abs/2312.15316 |