Spoken DialogSum: An Emotion-Rich Conversational Dataset for Spoken Dialogue Summarization

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Lu, Yen-Ju, Gao, Kunxiao, Liang, Mingrui, Wang, Helin, Thebaud, Thomas, Moro-Velazquez, Laureano, Dehak, Najim, Villalba, Jesus
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866917153102364672
author Lu, Yen-Ju
Gao, Kunxiao
Liang, Mingrui
Wang, Helin
Thebaud, Thomas
Moro-Velazquez, Laureano
Dehak, Najim
Villalba, Jesus
author_facet Lu, Yen-Ju
Gao, Kunxiao
Liang, Mingrui
Wang, Helin
Thebaud, Thomas
Moro-Velazquez, Laureano
Dehak, Najim
Villalba, Jesus
contents Recent audio language models can follow long conversations. However, research on emotion-aware or spoken dialogue summarization is constrained by the lack of data that links speech, summaries, and paralinguistic cues. We introduce Spoken DialogSum, the first corpus aligning raw conversational audio with factual summaries, emotion-rich summaries, and utterance-level labels for speaker age, gender, and emotion. The dataset is built in two stages: first, an LLM rewrites DialogSum scripts with Switchboard-style fillers and back-channels, then tags each utterance with emotion, pitch, and speaking rate. Second, an expressive TTS engine synthesizes speech from the tagged scripts, aligned with paralinguistic labels. Spoken DialogSum comprises 13,460 emotion-diverse dialogues, each paired with both a factual and an emotion-focused summary. We release an online demo at https://fatfat-emosum.github.io/EmoDialog-Sum-Audio-Samples/, with plans to release the full dataset in the near future. Baselines show that an Audio-LLM raises emotional-summary ROUGE-L by 28% relative to a cascaded ASR-LLM system, confirming the value of end-to-end speech modeling.
format Preprint
id arxiv_https___arxiv_org_abs_2512_14687
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Spoken DialogSum: An Emotion-Rich Conversational Dataset for Spoken Dialogue Summarization
Lu, Yen-Ju
Gao, Kunxiao
Liang, Mingrui
Wang, Helin
Thebaud, Thomas
Moro-Velazquez, Laureano
Dehak, Najim
Villalba, Jesus
Computation and Language
Artificial Intelligence
Machine Learning
Audio and Speech Processing
Recent audio language models can follow long conversations. However, research on emotion-aware or spoken dialogue summarization is constrained by the lack of data that links speech, summaries, and paralinguistic cues. We introduce Spoken DialogSum, the first corpus aligning raw conversational audio with factual summaries, emotion-rich summaries, and utterance-level labels for speaker age, gender, and emotion. The dataset is built in two stages: first, an LLM rewrites DialogSum scripts with Switchboard-style fillers and back-channels, then tags each utterance with emotion, pitch, and speaking rate. Second, an expressive TTS engine synthesizes speech from the tagged scripts, aligned with paralinguistic labels. Spoken DialogSum comprises 13,460 emotion-diverse dialogues, each paired with both a factual and an emotion-focused summary. We release an online demo at https://fatfat-emosum.github.io/EmoDialog-Sum-Audio-Samples/, with plans to release the full dataset in the near future. Baselines show that an Audio-LLM raises emotional-summary ROUGE-L by 28% relative to a cascaded ASR-LLM system, confirming the value of end-to-end speech modeling.
title Spoken DialogSum: An Emotion-Rich Conversational Dataset for Spoken Dialogue Summarization
topic Computation and Language
Artificial Intelligence
Machine Learning
Audio and Speech Processing
url https://arxiv.org/abs/2512.14687