HumDial-EIBench: A Human-Recorded Multi-Turn Emotional Intelligence Benchmark for Audio Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Shuiyuan, Zhao, Zhixian, Xue, Hongfei, Wang, Chengyou, Wang, Shuai, Bu, Hui, Xu, Xin, Xie, Lei
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914504420360192
author Wang, Shuiyuan
Zhao, Zhixian
Xue, Hongfei
Wang, Chengyou
Wang, Shuai
Bu, Hui
Xu, Xin
Xie, Lei
author_facet Wang, Shuiyuan
Zhao, Zhixian
Xue, Hongfei
Wang, Chengyou
Wang, Shuai
Bu, Hui
Xu, Xin
Xie, Lei
contents Evaluating the emotional intelligence (EI) of audio language models (ALMs) is critical. However, existing benchmarks mostly rely on synthesized speech, are limited to single-turn interactions, and depend heavily on open-ended scoring. This paper proposes HumDial-EIBench, a comprehensive benchmark for evaluating ALMs' EI. Using real-recorded human dialogues from the ICASSP 2026 HumDial Challenge, it reformulates emotional tracking and causal reasoning into multiple-choice questions with adversarial distractors, mitigating subjective scoring bias for cognitive tasks. It retains the generation of empathetic responses and introduces an acoustic-semantic conflict task to assess robustness against contradictory multimodal signals. Evaluations of eight ALMs reveal that most models struggle with multi-turn emotional tracking and implicit causal reasoning. Furthermore, all models exhibit decoupled textual and acoustic empathy, alongside a severe text-dominance bias during cross-modal conflicts.
format Preprint
id arxiv_https___arxiv_org_abs_2604_11594
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle HumDial-EIBench: A Human-Recorded Multi-Turn Emotional Intelligence Benchmark for Audio Language Models
Wang, Shuiyuan
Zhao, Zhixian
Xue, Hongfei
Wang, Chengyou
Wang, Shuai
Bu, Hui
Xu, Xin
Xie, Lei
Audio and Speech Processing
Sound
Evaluating the emotional intelligence (EI) of audio language models (ALMs) is critical. However, existing benchmarks mostly rely on synthesized speech, are limited to single-turn interactions, and depend heavily on open-ended scoring. This paper proposes HumDial-EIBench, a comprehensive benchmark for evaluating ALMs' EI. Using real-recorded human dialogues from the ICASSP 2026 HumDial Challenge, it reformulates emotional tracking and causal reasoning into multiple-choice questions with adversarial distractors, mitigating subjective scoring bias for cognitive tasks. It retains the generation of empathetic responses and introduces an acoustic-semantic conflict task to assess robustness against contradictory multimodal signals. Evaluations of eight ALMs reveal that most models struggle with multi-turn emotional tracking and implicit causal reasoning. Furthermore, all models exhibit decoupled textual and acoustic empathy, alongside a severe text-dominance bias during cross-modal conflicts.
title HumDial-EIBench: A Human-Recorded Multi-Turn Emotional Intelligence Benchmark for Audio Language Models
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2604.11594