EchoMind: An Interrelated Multi-level Benchmark for Evaluating Empathetic Speech Language Models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Zhou, Li, Yu, Lutong, Lyu, You, Lin, Yihang, Zhao, Zefeng, Ao, Junyi, Zhang, Yuhao, Wang, Benyou, Li, Haizhou
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866910041327534080
author Zhou, Li
Yu, Lutong
Lyu, You
Lin, Yihang
Zhao, Zefeng
Ao, Junyi
Zhang, Yuhao
Wang, Benyou
Li, Haizhou
author_facet Zhou, Li
Yu, Lutong
Lyu, You
Lin, Yihang
Zhao, Zefeng
Ao, Junyi
Zhang, Yuhao
Wang, Benyou
Li, Haizhou
contents Speech Language Models (SLMs) have made significant progress in spoken language understanding. Yet it remains unclear whether they can fully perceive non lexical vocal cues alongside spoken words, and respond with empathy that aligns with both emotional and contextual factors. Existing benchmarks typically evaluate linguistic, acoustic, reasoning, or dialogue abilities in isolation, overlooking the integration of these skills that is crucial for human-like, emotionally intelligent conversation. We present EchoMind, the first interrelated, multi-level benchmark that simulates the cognitive process of empathetic dialogue through sequential, context-linked tasks: spoken-content understanding, vocal-cue perception, integrated reasoning, and response generation. All tasks share identical and semantically neutral scripts that are free of explicit emotional or contextual cues, and controlled variations in vocal style are used to test the effect of delivery independent of the transcript. EchoMind is grounded in an empathy-oriented framework spanning 3 coarse and 12 fine-grained dimensions, encompassing 39 vocal attributes, and evaluated using both objective and subjective metrics. Testing 12 advanced SLMs reveals that even state-of-the-art models struggle with high-expressive vocal cues, limiting empathetic response quality. Analyses of prompt strength, speech source, and ideal vocal cue recognition reveal persistent weaknesses in instruction-following, resilience to natural speech variability, and effective use of vocal cues for empathy. These results underscore the need for SLMs that integrate linguistic content with diverse vocal cues to achieve truly empathetic conversational ability.
format Preprint
id arxiv_https___arxiv_org_abs_2510_22758
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle EchoMind: An Interrelated Multi-level Benchmark for Evaluating Empathetic Speech Language Models
Zhou, Li
Yu, Lutong
Lyu, You
Lin, Yihang
Zhao, Zefeng
Ao, Junyi
Zhang, Yuhao
Wang, Benyou
Li, Haizhou
Computation and Language
Speech Language Models (SLMs) have made significant progress in spoken language understanding. Yet it remains unclear whether they can fully perceive non lexical vocal cues alongside spoken words, and respond with empathy that aligns with both emotional and contextual factors. Existing benchmarks typically evaluate linguistic, acoustic, reasoning, or dialogue abilities in isolation, overlooking the integration of these skills that is crucial for human-like, emotionally intelligent conversation. We present EchoMind, the first interrelated, multi-level benchmark that simulates the cognitive process of empathetic dialogue through sequential, context-linked tasks: spoken-content understanding, vocal-cue perception, integrated reasoning, and response generation. All tasks share identical and semantically neutral scripts that are free of explicit emotional or contextual cues, and controlled variations in vocal style are used to test the effect of delivery independent of the transcript. EchoMind is grounded in an empathy-oriented framework spanning 3 coarse and 12 fine-grained dimensions, encompassing 39 vocal attributes, and evaluated using both objective and subjective metrics. Testing 12 advanced SLMs reveals that even state-of-the-art models struggle with high-expressive vocal cues, limiting empathetic response quality. Analyses of prompt strength, speech source, and ideal vocal cue recognition reveal persistent weaknesses in instruction-following, resilience to natural speech variability, and effective use of vocal cues for empathy. These results underscore the need for SLMs that integrate linguistic content with diverse vocal cues to achieve truly empathetic conversational ability.
title EchoMind: An Interrelated Multi-level Benchmark for Evaluating Empathetic Speech Language Models
topic Computation and Language
url https://arxiv.org/abs/2510.22758