EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Spoken Dialogue Systems
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866914005162917888 |
|---|---|
| author | Liu, Jingwen Cheng, Kan Jen Lian, Jiachen Anand, Akshay Jain, Rishi Qiao, Faith Netzorg, Robin Chou, Huang-Cheng Li, Tingle Lin, Guan-Ting Anumanchipalli, Gopala |
| author_facet | Liu, Jingwen Cheng, Kan Jen Lian, Jiachen Anand, Akshay Jain, Rishi Qiao, Faith Netzorg, Robin Chou, Huang-Cheng Li, Tingle Lin, Guan-Ting Anumanchipalli, Gopala |
| contents | Speech emotions play a crucial role in human-computer interaction, shaping engagement and context-aware communication. Despite recent advances in spoken dialogue systems, a holistic system for evaluating emotional reasoning is still lacking. To address this, we introduce EMO-Reasoning, a benchmark for assessing emotional coherence in dialogue systems. It leverages a curated dataset generated via text-to-speech to simulate diverse emotional states, overcoming the scarcity of emotional speech data. We further propose the Cross-turn Emotion Reasoning Score to assess the emotion transitions in multi-turn dialogues. Evaluating seven dialogue systems through continuous, categorical, and perceptual metrics, we show that our framework effectively detects emotional inconsistencies, providing insights for improving current dialogue systems. By releasing a systematic evaluation benchmark, we aim to advance emotion-aware spoken dialogue modeling toward more natural and adaptive interactions. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2508_17623 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Spoken Dialogue Systems Liu, Jingwen Cheng, Kan Jen Lian, Jiachen Anand, Akshay Jain, Rishi Qiao, Faith Netzorg, Robin Chou, Huang-Cheng Li, Tingle Lin, Guan-Ting Anumanchipalli, Gopala Computation and Language Audio and Speech Processing Speech emotions play a crucial role in human-computer interaction, shaping engagement and context-aware communication. Despite recent advances in spoken dialogue systems, a holistic system for evaluating emotional reasoning is still lacking. To address this, we introduce EMO-Reasoning, a benchmark for assessing emotional coherence in dialogue systems. It leverages a curated dataset generated via text-to-speech to simulate diverse emotional states, overcoming the scarcity of emotional speech data. We further propose the Cross-turn Emotion Reasoning Score to assess the emotion transitions in multi-turn dialogues. Evaluating seven dialogue systems through continuous, categorical, and perceptual metrics, we show that our framework effectively detects emotional inconsistencies, providing insights for improving current dialogue systems. By releasing a systematic evaluation benchmark, we aim to advance emotion-aware spoken dialogue modeling toward more natural and adaptive interactions. |
| title | EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Spoken Dialogue Systems |
| topic | Computation and Language Audio and Speech Processing |
| url | https://arxiv.org/abs/2508.17623 |