EmoBox: Multilingual Multi-corpus Speech Emotion Recognition Toolkit and Benchmark
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ma, Ziyang, Chen, Mingjie, Zhang, Hezhao, Zheng, Zhisheng, Chen, Wenxi, Li, Xiquan, Ye, Jiaxin, Chen, Xie, Hain, Thomas |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
VoxEmo: Benchmarking Speech Emotion Recognition with Speech LLMs
von: Zhang, Hezhao, et al.
Veröffentlicht: (2026)
von: Zhang, Hezhao, et al.
Veröffentlicht: (2026)
EmoDubber: Towards High Quality and Emotion Controllable Movie Dubbing
von: Cong, Gaoxiang, et al.
Veröffentlicht: (2024)
von: Cong, Gaoxiang, et al.
Veröffentlicht: (2024)
MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix
von: Ma, Ziyang, et al.
Veröffentlicht: (2025)
von: Ma, Ziyang, et al.
Veröffentlicht: (2025)
VERSA: A Versatile Evaluation Toolkit for Speech, Audio, and Music
von: Shi, Jiatong, et al.
Veröffentlicht: (2024)
von: Shi, Jiatong, et al.
Veröffentlicht: (2024)
A Survey on Multimodal Music Emotion Recognition
von: Liyanarachchi, Rashini, et al.
Veröffentlicht: (2025)
von: Liyanarachchi, Rashini, et al.
Veröffentlicht: (2025)
Multimodal Emotion Recognition from Raw Audio with Sinc-convolution
von: Zhang, Xiaohui, et al.
Veröffentlicht: (2024)
von: Zhang, Xiaohui, et al.
Veröffentlicht: (2024)
MSAC: Multiple Speech Attribute Control Method for Reliable Speech Emotion Recognition
von: Pan, Yu, et al.
Veröffentlicht: (2023)
von: Pan, Yu, et al.
Veröffentlicht: (2023)
Towards Reliable Large Audio Language Model
von: Ma, Ziyang, et al.
Veröffentlicht: (2025)
von: Ma, Ziyang, et al.
Veröffentlicht: (2025)
LCB-net: Long-Context Biasing for Audio-Visual Speech Recognition
von: Yu, Fan, et al.
Veröffentlicht: (2024)
von: Yu, Fan, et al.
Veröffentlicht: (2024)
Human-Inspired Computing for Robust and Efficient Audio-Visual Speech Recognition
von: Liu, Qianhui, et al.
Veröffentlicht: (2024)
von: Liu, Qianhui, et al.
Veröffentlicht: (2024)
DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis
von: Tian, Wenjie, et al.
Veröffentlicht: (2025)
von: Tian, Wenjie, et al.
Veröffentlicht: (2025)
SLAM-AAC: Enhancing Audio Captioning with Paraphrasing Augmentation and CLAP-Refine through LLMs
von: Chen, Wenxi, et al.
Veröffentlicht: (2024)
von: Chen, Wenxi, et al.
Veröffentlicht: (2024)
EmoReg: Directional Latent Vector Modeling for Emotional Intensity Regularization in Diffusion-based Voice Conversion
von: Gudmalwar, Ashishkumar, et al.
Veröffentlicht: (2024)
von: Gudmalwar, Ashishkumar, et al.
Veröffentlicht: (2024)
1st Place Solution to Odyssey Emotion Recognition Challenge Task1: Tackling Class Imbalance Problem
von: Chen, Mingjie, et al.
Veröffentlicht: (2024)
von: Chen, Mingjie, et al.
Veröffentlicht: (2024)
Robust LLM-based Audio-Visual Speech Recognition with Sparse Modality Alignment and Visual Unit-Guided Refinement
von: Su, Fei, et al.
Veröffentlicht: (2026)
von: Su, Fei, et al.
Veröffentlicht: (2026)
pTSE-T: Presentation Target Speaker Extraction using Unaligned Text Cues
von: Jiang, Ziyang, et al.
Veröffentlicht: (2024)
von: Jiang, Ziyang, et al.
Veröffentlicht: (2024)
pyAMPACT: A Score-Audio Alignment Toolkit for Performance Data Estimation and Multi-modal Processing
von: Devaney, Johanna, et al.
Veröffentlicht: (2024)
von: Devaney, Johanna, et al.
Veröffentlicht: (2024)
Sonic4D: Spatial Audio Generation for Immersive 4D Scene Exploration
von: Xie, Siyi, et al.
Veröffentlicht: (2025)
von: Xie, Siyi, et al.
Veröffentlicht: (2025)
Audio-CoT: Exploring Chain-of-Thought Reasoning in Large Audio Language Model
von: Ma, Ziyang, et al.
Veröffentlicht: (2025)
von: Ma, Ziyang, et al.
Veröffentlicht: (2025)
Automatic Speech Recognition System-Independent Word Error Rate Estimation
von: Park, Chanho, et al.
Veröffentlicht: (2024)
von: Park, Chanho, et al.
Veröffentlicht: (2024)
Speech Emotion Recognition with ASR Transcripts: A Comprehensive Study on Word Error Rate and Fusion Techniques
von: Li, Yuanchao, et al.
Veröffentlicht: (2024)
von: Li, Yuanchao, et al.
Veröffentlicht: (2024)
Stimulus Modality Matters: Impact of Perceptual Evaluations from Different Modalities on Speech Emotion Recognition System Performance
von: Chou, Huang-Cheng, et al.
Veröffentlicht: (2024)
von: Chou, Huang-Cheng, et al.
Veröffentlicht: (2024)
Intelligent Text-Conditioned Music Generation
von: Xie, Zhouyao, et al.
Veröffentlicht: (2024)
von: Xie, Zhouyao, et al.
Veröffentlicht: (2024)
Low-latency Speech Enhancement via Speech Token Generation
von: Xue, Huaying, et al.
Veröffentlicht: (2023)
von: Xue, Huaying, et al.
Veröffentlicht: (2023)
Efficient Speech Watermarking for Speech Synthesis via Progressive Knowledge Distillation
von: Cui, Yang, et al.
Veröffentlicht: (2025)
von: Cui, Yang, et al.
Veröffentlicht: (2025)
Speech-to-See: End-to-End Speech-Driven Open-Set Object Detection
von: Lu, Wenhuan, et al.
Veröffentlicht: (2025)
von: Lu, Wenhuan, et al.
Veröffentlicht: (2025)
MF-AED-AEC: Speech Emotion Recognition by Leveraging Multimodal Fusion, Asr Error Detection, and Asr Error Correction
von: He, Jiajun, et al.
Veröffentlicht: (2024)
von: He, Jiajun, et al.
Veröffentlicht: (2024)
Preserving Speaker Information in Direct Speech-to-Speech Translation with Non-Autoregressive Generation and Pretraining
von: Zhou, Rui, et al.
Veröffentlicht: (2024)
von: Zhou, Rui, et al.
Veröffentlicht: (2024)
Emotion-Aligned Contrastive Learning Between Images and Music
von: Stewart, Shanti, et al.
Veröffentlicht: (2023)
von: Stewart, Shanti, et al.
Veröffentlicht: (2023)
Conformer-based Ultrasound-to-Speech Conversion
von: Ibrahimov, Ibrahim, et al.
Veröffentlicht: (2025)
von: Ibrahimov, Ibrahim, et al.
Veröffentlicht: (2025)
Addressing Emotion Bias in Music Emotion Recognition and Generation with Frechet Audio Distance
von: Li, Yuanchao, et al.
Veröffentlicht: (2024)
von: Li, Yuanchao, et al.
Veröffentlicht: (2024)
Bimodal Connection Attention Fusion for Speech Emotion Recognition
von: Luo, Jiachen, et al.
Veröffentlicht: (2025)
von: Luo, Jiachen, et al.
Veröffentlicht: (2025)
Audio-Visual Speech Separation via Bottleneck Iterative Network
von: Zhang, Sidong, et al.
Veröffentlicht: (2025)
von: Zhang, Sidong, et al.
Veröffentlicht: (2025)
Zero-Shot Fake Video Detection by Audio-Visual Consistency
von: Li, Xiaolou, et al.
Veröffentlicht: (2024)
von: Li, Xiaolou, et al.
Veröffentlicht: (2024)
Emotion-Aware Speech Generation with Character-Specific Voices for Comics
von: Qian, Zhiwen, et al.
Veröffentlicht: (2025)
von: Qian, Zhiwen, et al.
Veröffentlicht: (2025)
Dopamine Audiobook: A Training-free MLLM Agent for Emotional and Immersive Audiobook Generation
von: Rong, Yan, et al.
Veröffentlicht: (2025)
von: Rong, Yan, et al.
Veröffentlicht: (2025)
Listening Between the Lines: Synthetic Speech Detection Disregarding Verbal Content
von: Salvi, Davide, et al.
Veröffentlicht: (2024)
von: Salvi, Davide, et al.
Veröffentlicht: (2024)
FastTalker: Jointly Generating Speech and Conversational Gestures from Text
von: Guo, Zixin, et al.
Veröffentlicht: (2024)
von: Guo, Zixin, et al.
Veröffentlicht: (2024)
Beyond Video-to-SFX: Video to Audio Synthesis with Environmentally Aware Speech
von: Niu, Xinlei, et al.
Veröffentlicht: (2025)
von: Niu, Xinlei, et al.
Veröffentlicht: (2025)
Leveraging Speech PTM, Text LLM, and Emotional TTS for Speech Emotion Recognition
von: Ma, Ziyang, et al.
Veröffentlicht: (2023)
von: Ma, Ziyang, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
VoxEmo: Benchmarking Speech Emotion Recognition with Speech LLMs
von: Zhang, Hezhao, et al.
Veröffentlicht: (2026) -
EmoDubber: Towards High Quality and Emotion Controllable Movie Dubbing
von: Cong, Gaoxiang, et al.
Veröffentlicht: (2024) -
MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix
von: Ma, Ziyang, et al.
Veröffentlicht: (2025) -
VERSA: A Versatile Evaluation Toolkit for Speech, Audio, and Music
von: Shi, Jiatong, et al.
Veröffentlicht: (2024) -
A Survey on Multimodal Music Emotion Recognition
von: Liyanarachchi, Rashini, et al.
Veröffentlicht: (2025)