EchoMind: An Interrelated Multi-level Benchmark for Evaluating Empathetic Speech Language Models
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Zhou, Li, Yu, Lutong, Lyu, You, Lin, Yihang, Zhao, Zefeng, Ao, Junyi, Zhang, Yuhao, Wang, Benyou, Li, Haizhou |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
EchoX: Towards Mitigating Acoustic-Semantic Gap via Echo Training for Speech-to-Speech LLMs
par: Zhang, Yuhao, et autres
Publié: (2025)
par: Zhang, Yuhao, et autres
Publié: (2025)
Roadmap towards Superhuman Speech Understanding using Large Language Models
par: Bu, Fan, et autres
Publié: (2024)
par: Bu, Fan, et autres
Publié: (2024)
MTalk-Bench: Evaluating Speech-to-Speech Models in Multi-Turn Dialogues via Arena-style and Rubrics Protocols
par: Du, Yuhao, et autres
Publié: (2025)
par: Du, Yuhao, et autres
Publié: (2025)
S2S-Arena: Evaluating Paralinguistic Instruction Following in Speech-to-Speech Models
par: Jiang, Feng, et autres
Publié: (2025)
par: Jiang, Feng, et autres
Publié: (2025)
Ti-Audio: The First Multi-Dialectal End-to-End Speech LLM for Tibetan
par: Wang, Jialing, et autres
Publié: (2026)
par: Wang, Jialing, et autres
Publié: (2026)
Soundwave: Less is More for Speech-Text Alignment in LLMs
par: Zhang, Yuhao, et autres
Publié: (2025)
par: Zhang, Yuhao, et autres
Publié: (2025)
SA-WavLM: Speaker-Aware Self-Supervised Pre-training for Mixture Speech
par: Lin, Jingru, et autres
Publié: (2024)
par: Lin, Jingru, et autres
Publié: (2024)
Chain-Talker: Chain Understanding and Rendering for Empathetic Conversational Speech Synthesis
par: Hu, Yifan, et autres
Publié: (2025)
par: Hu, Yifan, et autres
Publié: (2025)
From Word to World: Evaluate and Mitigate Culture Bias in LLMs via Word Association Test
par: Dai, Xunlian, et autres
Publié: (2025)
par: Dai, Xunlian, et autres
Publié: (2025)
Hanfu-Bench: A Multimodal Benchmark on Cross-Temporal Cultural Understanding and Transcreation
par: Zhou, Li, et autres
Publié: (2025)
par: Zhou, Li, et autres
Publié: (2025)
Text-guided HuBERT: Self-Supervised Speech Pre-training via Generative Adversarial Networks
par: Ma, Duo, et autres
Publié: (2024)
par: Ma, Duo, et autres
Publié: (2024)
Enhancing Empathetic Response Generation by Augmenting LLMs with Small-scale Empathetic Models
par: Yang, Zhou, et autres
Publié: (2024)
par: Yang, Zhou, et autres
Publié: (2024)
Leveraging Language Information for Target Language Extraction
par: Yıldırım, Mehmet Sinan, et autres
Publié: (2025)
par: Yıldırım, Mehmet Sinan, et autres
Publié: (2025)
Do We Really Need GNNs with Explicit Structural Modeling? MLPs Suffice for Language Model Representations
par: Zhou, Li, et autres
Publié: (2025)
par: Zhou, Li, et autres
Publié: (2025)
Beyond Binary: Towards Fine-Grained LLM-Generated Text Detection via Role Recognition and Involvement Measurement
par: Cheng, Zihao, et autres
Publié: (2024)
par: Cheng, Zihao, et autres
Publié: (2024)
An Iterative Associative Memory Model for Empathetic Response Generation
par: Yang, Zhou, et autres
Publié: (2024)
par: Yang, Zhou, et autres
Publié: (2024)
Multi-Scale Accent Modeling and Disentangling for Multi-Speaker Multi-Accent Text-to-Speech Synthesis
par: Zhou, Xuehao, et autres
Publié: (2024)
par: Zhou, Xuehao, et autres
Publié: (2024)
Computation and Parameter Efficient Multi-Modal Fusion Transformer for Cued Speech Recognition
par: Liu, Lei, et autres
Publié: (2024)
par: Liu, Lei, et autres
Publié: (2024)
Revolutionizing Database Q&A with Large Language Models: Comprehensive Benchmark and Evaluation
par: Zheng, Yihang, et autres
Publié: (2024)
par: Zheng, Yihang, et autres
Publié: (2024)
BLSP-Emo: Towards Empathetic Large Speech-Language Models
par: Wang, Chen, et autres
Publié: (2024)
par: Wang, Chen, et autres
Publié: (2024)
Multi-Step Prediction and Control of Hierarchical Emotion Distribution in Text-to-Speech Synthesis
par: Inoue, Sho, et autres
Publié: (2025)
par: Inoue, Sho, et autres
Publié: (2025)
Empathy Omni: Enabling Empathetic Speech Response Generation through Large Language Models
par: Wang, Haoyu, et autres
Publié: (2025)
par: Wang, Haoyu, et autres
Publié: (2025)
Multi-dimensional Evaluation of Empathetic Dialog Responses
par: Xu, Zhichao, et autres
Publié: (2024)
par: Xu, Zhichao, et autres
Publié: (2024)
Quantifying Self-diagnostic Atomic Knowledge in Chinese Medical Foundation Model: A Computational Analysis
par: Fan, Yaxin, et autres
Publié: (2023)
par: Fan, Yaxin, et autres
Publié: (2023)
Solla: Towards a Speech-Oriented LLM That Hears Acoustic Context
par: Ao, Junyi, et autres
Publié: (2025)
par: Ao, Junyi, et autres
Publié: (2025)
PERM: Psychology-grounded Empathetic Reward Modeling for Large Language Models
par: Wang, Chengbing, et autres
Publié: (2026)
par: Wang, Chengbing, et autres
Publié: (2026)
Leveraging Unit Language Guidance to Advance Speech Modeling in Textless Speech-to-Speech Translation
par: Zhang, Yuhao, et autres
Publié: (2025)
par: Zhang, Yuhao, et autres
Publié: (2025)
A Spectrum Evaluation Benchmark for Medical Multi-Modal Large Language Models
par: Liu, Jie, et autres
Publié: (2024)
par: Liu, Jie, et autres
Publié: (2024)
SD-Eval: A Benchmark Dataset for Spoken Dialogue Understanding Beyond Words
par: Ao, Junyi, et autres
Publié: (2024)
par: Ao, Junyi, et autres
Publié: (2024)
Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech
par: Liu, Rui, et autres
Publié: (2024)
par: Liu, Rui, et autres
Publié: (2024)
EDEN: Empathetic Dialogues for English learning
par: Siyan, Li, et autres
Publié: (2024)
par: Siyan, Li, et autres
Publié: (2024)
Emotionally Numb or Empathetic? Evaluating How LLMs Feel Using EmotionBench
par: Huang, Jen-tse, et autres
Publié: (2023)
par: Huang, Jen-tse, et autres
Publié: (2023)
Critical, Empathetic, and Mindful Relations (CEMR): A Relationship‐Building Theory
par: Sothy Eng
Publié: (2026)
par: Sothy Eng
Publié: (2026)
Hierarchical Control of Emotion Rendering in Speech Synthesis
par: Inoue, Sho, et autres
Publié: (2024)
par: Inoue, Sho, et autres
Publié: (2024)
Hierarchical Emotion Prediction and Control in Text-to-Speech Synthesis
par: Inoue, Sho, et autres
Publié: (2024)
par: Inoue, Sho, et autres
Publié: (2024)
Fine-Grained Quantitative Emotion Editing for Speech Generation
par: Inoue, Sho, et autres
Publié: (2024)
par: Inoue, Sho, et autres
Publié: (2024)
Rethinking Meeting Effectiveness: A Benchmark and Framework for Temporal Fine-grained Automatic Meeting Effectiveness Evaluation
par: Li, Yihang, et autres
Publié: (2026)
par: Li, Yihang, et autres
Publié: (2026)
Human or Machine? A Preliminary Turing Test for Speech-to-Speech Interaction
par: Li, Xiang, et autres
Publié: (2026)
par: Li, Xiang, et autres
Publié: (2026)
Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark
par: Zhang, Han, et autres
Publié: (2025)
par: Zhang, Han, et autres
Publié: (2025)
Multi-source Multi-level Multi-token Ethereum Dataset and Benchmark Platform
par: Li, Haoyuan, et autres
Publié: (2025)
par: Li, Haoyuan, et autres
Publié: (2025)
Documents similaires
-
EchoX: Towards Mitigating Acoustic-Semantic Gap via Echo Training for Speech-to-Speech LLMs
par: Zhang, Yuhao, et autres
Publié: (2025) -
Roadmap towards Superhuman Speech Understanding using Large Language Models
par: Bu, Fan, et autres
Publié: (2024) -
MTalk-Bench: Evaluating Speech-to-Speech Models in Multi-Turn Dialogues via Arena-style and Rubrics Protocols
par: Du, Yuhao, et autres
Publié: (2025) -
S2S-Arena: Evaluating Paralinguistic Instruction Following in Speech-to-Speech Models
par: Jiang, Feng, et autres
Publié: (2025) -
Ti-Audio: The First Multi-Dialectal End-to-End Speech LLM for Tibetan
par: Wang, Jialing, et autres
Publié: (2026)