EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Spoken Dialogue Systems
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Jingwen, Cheng, Kan Jen, Lian, Jiachen, Anand, Akshay, Jain, Rishi, Qiao, Faith, Netzorg, Robin, Chou, Huang-Cheng, Li, Tingle, Lin, Guan-Ting, Anumanchipalli, Gopala |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues
by: Zhou, Dingkun, et al.
Published: (2025)
by: Zhou, Dingkun, et al.
Published: (2025)
Audio Texture Manipulation by Exemplar-Based Analogy
by: Cheng, Kan Jen, et al.
Published: (2025)
by: Cheng, Kan Jen, et al.
Published: (2025)
Towards Hierarchical Spoken Language Dysfluency Modeling
by: Lian, Jiachen, et al.
Published: (2024)
by: Lian, Jiachen, et al.
Published: (2024)
Full-Duplex-Bench: A Benchmark to Evaluate Full-duplex Spoken Dialogue Models on Turn-taking Capabilities
by: Lin, Guan-Ting, et al.
Published: (2025)
by: Lin, Guan-Ting, et al.
Published: (2025)
Enabling Conversational Behavior Reasoning Capabilities in Full-Duplex Speech
by: Pan, Shuchang, et al.
Published: (2025)
by: Pan, Shuchang, et al.
Published: (2025)
Teaching Machines to Speak Using Articulatory Control
by: Anand, Akshay, et al.
Published: (2025)
by: Anand, Akshay, et al.
Published: (2025)
Speech World Model: Causal State-Action Planning with Explicit Reasoning for Speech
by: Zhou, Xuanru, et al.
Published: (2025)
by: Zhou, Xuanru, et al.
Published: (2025)
The Oracle Has Spoken: A Multi-Aspect Evaluation of Dialogue in Pythia
by: Chen, Zixun, et al.
Published: (2025)
by: Chen, Zixun, et al.
Published: (2025)
TART: A Comprehensive Tool for Technique-Aware Audio-to-Tab Guitar Transcription
by: Gupta, Akshaj, et al.
Published: (2025)
by: Gupta, Akshaj, et al.
Published: (2025)
Speech After Gender: A Trans-Feminine Perspective on Next Steps for Speech Science and Technology
by: Netzorg, Robin, et al.
Published: (2024)
by: Netzorg, Robin, et al.
Published: (2024)
Unsupervised TTS Acoustic Modeling for TTS with Conditional Disentangled Sequential VAE
by: Lian, Jiachen, et al.
Published: (2022)
by: Lian, Jiachen, et al.
Published: (2022)
Multimodal Segmentation for Vocal Tract Modeling
by: Jain, Rishi, et al.
Published: (2024)
by: Jain, Rishi, et al.
Published: (2024)
Self-Supervised Audio-Visual Soundscape Stylization
by: Li, Tingle, et al.
Published: (2024)
by: Li, Tingle, et al.
Published: (2024)
Scaling Spoken Language Models with Syllabic Speech Tokenization
by: Lee, Nicholas, et al.
Published: (2025)
by: Lee, Nicholas, et al.
Published: (2025)
EMO-Debias: Benchmarking Gender Debiasing Techniques in Multi-Label Speech Emotion Recognition
by: Lin, Yi-Cheng, et al.
Published: (2025)
by: Lin, Yi-Cheng, et al.
Published: (2025)
Asymmetric Hierarchical Anchoring for Audio-Visual Joint Representation: Resolving Information Allocation Ambiguity for Robust Cross-Modal Generalization
by: Wu, Bixing, et al.
Published: (2026)
by: Wu, Bixing, et al.
Published: (2026)
EMO-SUPERB: An In-depth Look at Speech Emotion Recognition
by: Wu, Haibin, et al.
Published: (2024)
by: Wu, Haibin, et al.
Published: (2024)
EMO-Codec: An In-Depth Look at Emotion Preservation capacity of Legacy and Neural Codec Models With Subjective and Objective Evaluations
by: Ren, Wenze, et al.
Published: (2024)
by: Ren, Wenze, et al.
Published: (2024)
A Unified Framework for Model Editing
by: Gupta, Akshat, et al.
Published: (2024)
by: Gupta, Akshat, et al.
Published: (2024)
Is Bigger Edit Batch Size Always Better? -- An Empirical Study on Model Editing with Llama-3
by: Yoon, Junsang, et al.
Published: (2024)
by: Yoon, Junsang, et al.
Published: (2024)
Rebuilding ROME : Resolving Model Collapse during Sequential Model Editing
by: Gupta, Akshat, et al.
Published: (2024)
by: Gupta, Akshat, et al.
Published: (2024)
Geometric Interpretation of Layer Normalization and a Comparative Analysis with RMSNorm
by: Gupta, Akshat, et al.
Published: (2024)
by: Gupta, Akshat, et al.
Published: (2024)
Model Editing at Scale leads to Gradual and Catastrophic Forgetting
by: Gupta, Akshat, et al.
Published: (2024)
by: Gupta, Akshat, et al.
Published: (2024)
StyleStream: Real-Time Zero-Shot Voice Style Conversion
by: Liu, Yisi, et al.
Published: (2026)
by: Liu, Yisi, et al.
Published: (2026)
Self-Assessment Tests are Unreliable Measures of LLM Personality
by: Gupta, Akshat, et al.
Published: (2023)
by: Gupta, Akshat, et al.
Published: (2023)
HuPER: A Human-Inspired Framework for Phonetic Perception
by: Guo, Chenxu, et al.
Published: (2026)
by: Guo, Chenxu, et al.
Published: (2026)
EMO-R3: Reflective Reinforcement Learning for Emotional Reasoning in Multimodal Large Language Models
by: Fang, Yiyang, et al.
Published: (2026)
by: Fang, Yiyang, et al.
Published: (2026)
Schrodinger Audio-Visual Editor: Object-Level Audiovisual Removal
by: Xu, Weihan, et al.
Published: (2025)
by: Xu, Weihan, et al.
Published: (2025)
EmoTrans: A Benchmark for Understanding, Reasoning, and Predicting Emotion Transitions in Multimodal LLMs
by: Hu, He, et al.
Published: (2026)
by: Hu, He, et al.
Published: (2026)
Spoken DialogSum: An Emotion-Rich Conversational Dataset for Spoken Dialogue Summarization
by: Lu, Yen-Ju, et al.
Published: (2025)
by: Lu, Yen-Ju, et al.
Published: (2025)
Sounding that Object: Interactive Object-Aware Image to Audio Generation
by: Li, Tingle, et al.
Published: (2025)
by: Li, Tingle, et al.
Published: (2025)
DeepDialogue: A Multi-Turn Emotionally-Rich Spoken Dialogue Dataset
by: Koudounas, Alkis, et al.
Published: (2025)
by: Koudounas, Alkis, et al.
Published: (2025)
Emotion-LLaMA: Multimodal Emotion Recognition and Reasoning with Instruction Tuning
by: Cheng, Zebang, et al.
Published: (2024)
by: Cheng, Zebang, et al.
Published: (2024)
Improving Prototypical Visual Explanations with Reward Reweighing, Reselection, and Retraining
by: Li, Aaron J., et al.
Published: (2023)
by: Li, Aaron J., et al.
Published: (2023)
nEMO: Dataset of Emotional Speech in Polish
by: Christop, Iwona
Published: (2024)
by: Christop, Iwona
Published: (2024)
How Do LLMs Use Their Depth?
by: Gupta, Akshat, et al.
Published: (2025)
by: Gupta, Akshat, et al.
Published: (2025)
Efficient Knowledge Editing via Minimal Precomputation
by: Gupta, Akshat, et al.
Published: (2025)
by: Gupta, Akshat, et al.
Published: (2025)
Chain-of-Thought Reasoning in Streaming Full-Duplex End-to-End Spoken Dialogue Systems
by: Arora, Siddhant, et al.
Published: (2025)
by: Arora, Siddhant, et al.
Published: (2025)
WavBench: Benchmarking Reasoning, Colloquialism, and Paralinguistics for End-to-End Spoken Dialogue Models
by: Li, Yangzhuo, et al.
Published: (2026)
by: Li, Yangzhuo, et al.
Published: (2026)
EMO-KNOW: A Large Scale Dataset on Emotion and Emotion-cause
by: Nguyen, Mia Huong, et al.
Published: (2024)
by: Nguyen, Mia Huong, et al.
Published: (2024)
Similar Items
-
AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues
by: Zhou, Dingkun, et al.
Published: (2025) -
Audio Texture Manipulation by Exemplar-Based Analogy
by: Cheng, Kan Jen, et al.
Published: (2025) -
Towards Hierarchical Spoken Language Dysfluency Modeling
by: Lian, Jiachen, et al.
Published: (2024) -
Full-Duplex-Bench: A Benchmark to Evaluate Full-duplex Spoken Dialogue Models on Turn-taking Capabilities
by: Lin, Guan-Ting, et al.
Published: (2025) -
Enabling Conversational Behavior Reasoning Capabilities in Full-Duplex Speech
by: Pan, Shuchang, et al.
Published: (2025)