Spoken DialogSum: An Emotion-Rich Conversational Dataset for Spoken Dialogue Summarization
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866917153102364672 |
|---|---|
| author | Lu, Yen-Ju Gao, Kunxiao Liang, Mingrui Wang, Helin Thebaud, Thomas Moro-Velazquez, Laureano Dehak, Najim Villalba, Jesus |
| author_facet | Lu, Yen-Ju Gao, Kunxiao Liang, Mingrui Wang, Helin Thebaud, Thomas Moro-Velazquez, Laureano Dehak, Najim Villalba, Jesus |
| contents | Recent audio language models can follow long conversations. However, research on emotion-aware or spoken dialogue summarization is constrained by the lack of data that links speech, summaries, and paralinguistic cues. We introduce Spoken DialogSum, the first corpus aligning raw conversational audio with factual summaries, emotion-rich summaries, and utterance-level labels for speaker age, gender, and emotion. The dataset is built in two stages: first, an LLM rewrites DialogSum scripts with Switchboard-style fillers and back-channels, then tags each utterance with emotion, pitch, and speaking rate. Second, an expressive TTS engine synthesizes speech from the tagged scripts, aligned with paralinguistic labels. Spoken DialogSum comprises 13,460 emotion-diverse dialogues, each paired with both a factual and an emotion-focused summary. We release an online demo at https://fatfat-emosum.github.io/EmoDialog-Sum-Audio-Samples/, with plans to release the full dataset in the near future. Baselines show that an Audio-LLM raises emotional-summary ROUGE-L by 28% relative to a cascaded ASR-LLM system, confirming the value of end-to-end speech modeling. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2512_14687 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Spoken DialogSum: An Emotion-Rich Conversational Dataset for Spoken Dialogue Summarization Lu, Yen-Ju Gao, Kunxiao Liang, Mingrui Wang, Helin Thebaud, Thomas Moro-Velazquez, Laureano Dehak, Najim Villalba, Jesus Computation and Language Artificial Intelligence Machine Learning Audio and Speech Processing Recent audio language models can follow long conversations. However, research on emotion-aware or spoken dialogue summarization is constrained by the lack of data that links speech, summaries, and paralinguistic cues. We introduce Spoken DialogSum, the first corpus aligning raw conversational audio with factual summaries, emotion-rich summaries, and utterance-level labels for speaker age, gender, and emotion. The dataset is built in two stages: first, an LLM rewrites DialogSum scripts with Switchboard-style fillers and back-channels, then tags each utterance with emotion, pitch, and speaking rate. Second, an expressive TTS engine synthesizes speech from the tagged scripts, aligned with paralinguistic labels. Spoken DialogSum comprises 13,460 emotion-diverse dialogues, each paired with both a factual and an emotion-focused summary. We release an online demo at https://fatfat-emosum.github.io/EmoDialog-Sum-Audio-Samples/, with plans to release the full dataset in the near future. Baselines show that an Audio-LLM raises emotional-summary ROUGE-L by 28% relative to a cascaded ASR-LLM system, confirming the value of end-to-end speech modeling. |
| title | Spoken DialogSum: An Emotion-Rich Conversational Dataset for Spoken Dialogue Summarization |
| topic | Computation and Language Artificial Intelligence Machine Learning Audio and Speech Processing |
| url | https://arxiv.org/abs/2512.14687 |