EMO-SUPERB: An In-depth Look at Speech Emotion Recognition
Fuente:
arXiv
Salvato in:
| Autori principali: | Wu, Haibin, Chou, Huang-Cheng, Chang, Kai-Wei, Goncalves, Lucas, Du, Jiawei, Jang, Jyh-Shing Roger, Lee, Chi-Chun, Lee, Hung-Yi |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
EMO-Codec: An In-Depth Look at Emotion Preservation capacity of Legacy and Neural Codec Models With Subjective and Objective Evaluations
di: Ren, Wenze, et al.
Pubblicazione: (2024)
di: Ren, Wenze, et al.
Pubblicazione: (2024)
Neural Codec-based Adversarial Sample Detection for Speaker Verification
di: Chen, Xuanjun, et al.
Pubblicazione: (2024)
di: Chen, Xuanjun, et al.
Pubblicazione: (2024)
Stimulus Modality Matters: Impact of Perceptual Evaluations from Different Modalities on Speech Emotion Recognition System Performance
di: Chou, Huang-Cheng, et al.
Pubblicazione: (2024)
di: Chou, Huang-Cheng, et al.
Pubblicazione: (2024)
Codec-Based Deepfake Source Tracing via Neural Audio Codec Taxonomy
di: Chen, Xuanjun, et al.
Pubblicazione: (2025)
di: Chen, Xuanjun, et al.
Pubblicazione: (2025)
Singing Voice Graph Modeling for SingFake Detection
di: Chen, Xuanjun, et al.
Pubblicazione: (2024)
di: Chen, Xuanjun, et al.
Pubblicazione: (2024)
Codec-SUPERB: An In-Depth Analysis of Sound Codec Models
di: Wu, Haibin, et al.
Pubblicazione: (2024)
di: Wu, Haibin, et al.
Pubblicazione: (2024)
CodecFake+: A Large-Scale Neural Audio Codec-Based Deepfake Speech Dataset
di: Chen, Xuanjun, et al.
Pubblicazione: (2025)
di: Chen, Xuanjun, et al.
Pubblicazione: (2025)
Dynamic-SUPERB: Towards A Dynamic, Collaborative, and Comprehensive Instruction-Tuning Benchmark for Speech
di: Huang, Chien-yu, et al.
Pubblicazione: (2023)
di: Huang, Chien-yu, et al.
Pubblicazione: (2023)
DFADD: The Diffusion and Flow-Matching Based Audio Deepfake Dataset
di: Du, Jiawei, et al.
Pubblicazione: (2024)
di: Du, Jiawei, et al.
Pubblicazione: (2024)
EMO-RL: Emotion-Rule-Based Reinforcement Learning Enhanced Audio-Language Model for Generalized Speech Emotion Recognition
di: Li, Pengcheng, et al.
Pubblicazione: (2025)
di: Li, Pengcheng, et al.
Pubblicazione: (2025)
Towards Generalized Source Tracing for Codec-Based Deepfake Speech
di: Chen, Xuanjun, et al.
Pubblicazione: (2025)
di: Chen, Xuanjun, et al.
Pubblicazione: (2025)
Codec-SUPERB @ SLT 2024: A lightweight benchmark for neural audio codec models
di: Wu, Haibin, et al.
Pubblicazione: (2024)
di: Wu, Haibin, et al.
Pubblicazione: (2024)
nEMO: Dataset of Emotional Speech in Polish
di: Christop, Iwona
Pubblicazione: (2024)
di: Christop, Iwona
Pubblicazione: (2024)
Improving Speech Emotion Recognition in Under-Resourced Languages via Speech-to-Speech Translation with Bootstrapping Data Selection
di: Lin, Hsi-Che, et al.
Pubblicazione: (2024)
di: Lin, Hsi-Che, et al.
Pubblicazione: (2024)
ML-SUPERB: Multilingual Speech Universal PERformance Benchmark
di: Shi, Jiatong, et al.
Pubblicazione: (2023)
di: Shi, Jiatong, et al.
Pubblicazione: (2023)
Mitigating Subgroup Disparities in Multi-Label Speech Emotion Recognition: A Pseudo-Labeling and Unsupervised Learning Approach
di: Lin, Yi-Cheng, et al.
Pubblicazione: (2025)
di: Lin, Yi-Cheng, et al.
Pubblicazione: (2025)
A Layer-Anchoring Strategy for Enhancing Cross-Lingual Speech Emotion Recognition
di: Upadhyay, Shreya G., et al.
Pubblicazione: (2024)
di: Upadhyay, Shreya G., et al.
Pubblicazione: (2024)
Joint Fullband-Subband Modeling for High-Resolution SingFake Detection
di: Chen, Xuanjun, et al.
Pubblicazione: (2026)
di: Chen, Xuanjun, et al.
Pubblicazione: (2026)
CodecFake: Enhancing Anti-Spoofing Models Against Deepfake Audios from Codec-Based Speech Synthesis Systems
di: Wu, Haibin, et al.
Pubblicazione: (2024)
di: Wu, Haibin, et al.
Pubblicazione: (2024)
Dataset-Distillation Generative Model for Speech Emotion Recognition
di: Ritter-Gutierrez, Fabian, et al.
Pubblicazione: (2024)
di: Ritter-Gutierrez, Fabian, et al.
Pubblicazione: (2024)
Mouth Articulation-Based Anchoring for Improved Cross-Corpus Speech Emotion Recognition
di: Upadhyay, Shreya G., et al.
Pubblicazione: (2024)
di: Upadhyay, Shreya G., et al.
Pubblicazione: (2024)
Emo-bias: A Large Scale Evaluation of Social Bias on Speech Emotion Recognition
di: Lin, Yi-Cheng, et al.
Pubblicazione: (2024)
di: Lin, Yi-Cheng, et al.
Pubblicazione: (2024)
Continual Test-time Adaptation for End-to-end Speech Recognition on Noisy Speech
di: Lin, Guan-Ting, et al.
Pubblicazione: (2024)
di: Lin, Guan-Ting, et al.
Pubblicazione: (2024)
ML-SUPERB 2.0: Benchmarking Multilingual Speech Models Across Modeling Constraints, Languages, and Datasets
di: Shi, Jiatong, et al.
Pubblicazione: (2024)
di: Shi, Jiatong, et al.
Pubblicazione: (2024)
Singer separation for karaoke content generation
di: Lin, Hsuan-Yu, et al.
Pubblicazione: (2021)
di: Lin, Hsuan-Yu, et al.
Pubblicazione: (2021)
Unifying Listener Scoring Scales: Comparison Learning Framework for Speech Quality Assessment and Continuous Speech Emotion Recognition
di: Hu, Cheng-Hung, et al.
Pubblicazione: (2025)
di: Hu, Cheng-Hung, et al.
Pubblicazione: (2025)
Token-Level Logits Matter: A Closer Look at Speech Foundation Models for Ambiguous Emotion Recognition
di: Halim, Jule Valendo, et al.
Pubblicazione: (2025)
di: Halim, Jule Valendo, et al.
Pubblicazione: (2025)
Multimodal Transformer Distillation for Audio-Visual Synchronization
di: Chen, Xuanjun, et al.
Pubblicazione: (2022)
di: Chen, Xuanjun, et al.
Pubblicazione: (2022)
Speech Emotion Recognition with ASR Integration
di: Li, Yuanchao
Pubblicazione: (2026)
di: Li, Yuanchao
Pubblicazione: (2026)
EMO-Debias: Benchmarking Gender Debiasing Techniques in Multi-Label Speech Emotion Recognition
di: Lin, Yi-Cheng, et al.
Pubblicazione: (2025)
di: Lin, Yi-Cheng, et al.
Pubblicazione: (2025)
Emotion Neural Transducer for Fine-Grained Speech Emotion Recognition
di: Shen, Siyuan, et al.
Pubblicazione: (2024)
di: Shen, Siyuan, et al.
Pubblicazione: (2024)
Revisiting Modeling and Evaluation Approaches in Speech Emotion Recognition: Considering Subjectivity of Annotators and Ambiguity of Emotions
di: Chou, Huang-Cheng, et al.
Pubblicazione: (2025)
di: Chou, Huang-Cheng, et al.
Pubblicazione: (2025)
Crab: Multi Layer Contrastive Supervision to Improve Speech Emotion Recognition Under Both Acted and Natural Speech Condition
di: Ueda, Lucas H., et al.
Pubblicazione: (2026)
di: Ueda, Lucas H., et al.
Pubblicazione: (2026)
Improving Speech Emotion Recognition Through Cross Modal Attention Alignment and Balanced Stacking Model
di: Ueda, Lucas, et al.
Pubblicazione: (2025)
di: Ueda, Lucas, et al.
Pubblicazione: (2025)
SMILE: Speech Meta In-Context Learning for Low-Resource Language Automatic Speech Recognition
di: Hsu, Ming-Hao, et al.
Pubblicazione: (2024)
di: Hsu, Ming-Hao, et al.
Pubblicazione: (2024)
Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition
di: Wang, Shih-heng, et al.
Pubblicazione: (2024)
di: Wang, Shih-heng, et al.
Pubblicazione: (2024)
Toward Fair Speech Technologies: A Comprehensive Survey of Bias and Fairness in Speech AI
di: Lin, Yi-Cheng, et al.
Pubblicazione: (2026)
di: Lin, Yi-Cheng, et al.
Pubblicazione: (2026)
TS-SUPERB: A Target Speech Processing Benchmark for Speech Self-Supervised Learning Models
di: Peng, Junyi, et al.
Pubblicazione: (2025)
di: Peng, Junyi, et al.
Pubblicazione: (2025)
Toward Efficient Speech Emotion Recognition via Spectral Learning and Attention
di: Lee, HyeYoung, et al.
Pubblicazione: (2025)
di: Lee, HyeYoung, et al.
Pubblicazione: (2025)
THAI Speech Emotion Recognition (THAI-SER) corpus
di: Wongpithayadisai, Jilamika, et al.
Pubblicazione: (2025)
di: Wongpithayadisai, Jilamika, et al.
Pubblicazione: (2025)
Documenti analoghi
-
EMO-Codec: An In-Depth Look at Emotion Preservation capacity of Legacy and Neural Codec Models With Subjective and Objective Evaluations
di: Ren, Wenze, et al.
Pubblicazione: (2024) -
Neural Codec-based Adversarial Sample Detection for Speaker Verification
di: Chen, Xuanjun, et al.
Pubblicazione: (2024) -
Stimulus Modality Matters: Impact of Perceptual Evaluations from Different Modalities on Speech Emotion Recognition System Performance
di: Chou, Huang-Cheng, et al.
Pubblicazione: (2024) -
Codec-Based Deepfake Source Tracing via Neural Audio Codec Taxonomy
di: Chen, Xuanjun, et al.
Pubblicazione: (2025) -
Singing Voice Graph Modeling for SingFake Detection
di: Chen, Xuanjun, et al.
Pubblicazione: (2024)