Speech Emotion Recognition with ASR Transcripts: A Comprehensive Study on Word Error Rate and Fusion Techniques
Fuente:
arXiv
Salvato in:
| Autori principali: | Li, Yuanchao, Bell, Peter, Lai, Catherine |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Revise, Reason, and Recognize: LLM-Based Emotion Recognition via Emotion-Specific Prompts and ASR Error Correction
di: Li, Yuanchao, et al.
Pubblicazione: (2024)
di: Li, Yuanchao, et al.
Pubblicazione: (2024)
Crossmodal ASR Error Correction with Discrete Speech Units
di: Li, Yuanchao, et al.
Pubblicazione: (2024)
di: Li, Yuanchao, et al.
Pubblicazione: (2024)
MF-AED-AEC: Speech Emotion Recognition by Leveraging Multimodal Fusion, Asr Error Detection, and Asr Error Correction
di: He, Jiajun, et al.
Pubblicazione: (2024)
di: He, Jiajun, et al.
Pubblicazione: (2024)
Semi-Supervised Cognitive State Classification from Speech with Multi-View Pseudo-Labeling
di: Li, Yuanchao, et al.
Pubblicazione: (2024)
di: Li, Yuanchao, et al.
Pubblicazione: (2024)
Addressing Emotion Bias in Music Emotion Recognition and Generation with Frechet Audio Distance
di: Li, Yuanchao, et al.
Pubblicazione: (2024)
di: Li, Yuanchao, et al.
Pubblicazione: (2024)
Layer-Wise Analysis of Self-Supervised Acoustic Word Embeddings: A Study on Speech Emotion Recognition
di: Saliba, Alexandra, et al.
Pubblicazione: (2024)
di: Saliba, Alexandra, et al.
Pubblicazione: (2024)
Bimodal Connection Attention Fusion for Speech Emotion Recognition
di: Luo, Jiachen, et al.
Pubblicazione: (2025)
di: Luo, Jiachen, et al.
Pubblicazione: (2025)
Exploring Acoustic Similarity in Emotional Speech and Music via Self-Supervised Representations
di: Sun, Yujia, et al.
Pubblicazione: (2024)
di: Sun, Yujia, et al.
Pubblicazione: (2024)
Speech Emotion Recognition with ASR Integration
di: Li, Yuanchao
Pubblicazione: (2026)
di: Li, Yuanchao
Pubblicazione: (2026)
VoxEmo: Benchmarking Speech Emotion Recognition with Speech LLMs
di: Zhang, Hezhao, et al.
Pubblicazione: (2026)
di: Zhang, Hezhao, et al.
Pubblicazione: (2026)
MLLM-based Speech Recognition: When and How is Multimodality Beneficial?
di: Guan, Yiwen, et al.
Pubblicazione: (2025)
di: Guan, Yiwen, et al.
Pubblicazione: (2025)
EmoBox: Multilingual Multi-corpus Speech Emotion Recognition Toolkit and Benchmark
di: Ma, Ziyang, et al.
Pubblicazione: (2024)
di: Ma, Ziyang, et al.
Pubblicazione: (2024)
Mega-ASR: Towards In-the-wild^2 Speech Recognition via Scaling up Real-world Acoustic Simulation
di: Xie, Zhifei, et al.
Pubblicazione: (2026)
di: Xie, Zhifei, et al.
Pubblicazione: (2026)
A Comprehensive Study on the Effectiveness of ASR Representations for Noise-Robust Speech Emotion Recognition
di: Shi, Xiaohan, et al.
Pubblicazione: (2023)
di: Shi, Xiaohan, et al.
Pubblicazione: (2023)
Fretting-Transformer: Encoder-Decoder Model for MIDI to Tablature Transcription
di: Hamberger, Anna, et al.
Pubblicazione: (2025)
di: Hamberger, Anna, et al.
Pubblicazione: (2025)
Whispering LLaMA: A Cross-Modal Generative Error Correction Framework for Speech Recognition
di: Radhakrishnan, Srijith, et al.
Pubblicazione: (2023)
di: Radhakrishnan, Srijith, et al.
Pubblicazione: (2023)
ELEGANCE: Efficient LLM Guidance for Audio-Visual Target Speech Extraction
di: Wu, Wenxuan, et al.
Pubblicazione: (2025)
di: Wu, Wenxuan, et al.
Pubblicazione: (2025)
Double Mixture: Towards Continual Event Detection from Speech
di: Kang, Jingqi, et al.
Pubblicazione: (2024)
di: Kang, Jingqi, et al.
Pubblicazione: (2024)
MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix
di: Ma, Ziyang, et al.
Pubblicazione: (2025)
di: Ma, Ziyang, et al.
Pubblicazione: (2025)
CommonVoice-SpeechRE and RPG-MoGe: Advancing Speech Relation Extraction with a New Dataset and Multi-Order Generative Framework
di: Ning, Jinzhong, et al.
Pubblicazione: (2025)
di: Ning, Jinzhong, et al.
Pubblicazione: (2025)
StarVC: A Unified Auto-Regressive Framework for Joint Text and Speech Generation in Voice Conversion
di: Li, Fengjin, et al.
Pubblicazione: (2025)
di: Li, Fengjin, et al.
Pubblicazione: (2025)
Automatic Speech Recognition System-Independent Word Error Rate Estimation
di: Park, Chanho, et al.
Pubblicazione: (2024)
di: Park, Chanho, et al.
Pubblicazione: (2024)
Real-Time Word-Level Temporal Segmentation in Streaming Speech Recognition
di: Nishida, Naoto, et al.
Pubblicazione: (2025)
di: Nishida, Naoto, et al.
Pubblicazione: (2025)
A Survey on Multimodal Music Emotion Recognition
di: Liyanarachchi, Rashini, et al.
Pubblicazione: (2025)
di: Liyanarachchi, Rashini, et al.
Pubblicazione: (2025)
Human-Inspired Computing for Robust and Efficient Audio-Visual Speech Recognition
di: Liu, Qianhui, et al.
Pubblicazione: (2024)
di: Liu, Qianhui, et al.
Pubblicazione: (2024)
It's Never Too Late: Fusing Acoustic Information into Large Language Models for Automatic Speech Recognition
di: Chen, Chen, et al.
Pubblicazione: (2024)
di: Chen, Chen, et al.
Pubblicazione: (2024)
Multimodal Emotion Recognition from Raw Audio with Sinc-convolution
di: Zhang, Xiaohui, et al.
Pubblicazione: (2024)
di: Zhang, Xiaohui, et al.
Pubblicazione: (2024)
An Embarrassingly Simple Approach for LLM with Strong ASR Capacity
di: Ma, Ziyang, et al.
Pubblicazione: (2024)
di: Ma, Ziyang, et al.
Pubblicazione: (2024)
MSAC: Multiple Speech Attribute Control Method for Reliable Speech Emotion Recognition
di: Pan, Yu, et al.
Pubblicazione: (2023)
di: Pan, Yu, et al.
Pubblicazione: (2023)
LCB-net: Long-Context Biasing for Audio-Visual Speech Recognition
di: Yu, Fan, et al.
Pubblicazione: (2024)
di: Yu, Fan, et al.
Pubblicazione: (2024)
Robust LLM-based Audio-Visual Speech Recognition with Sparse Modality Alignment and Visual Unit-Guided Refinement
di: Su, Fei, et al.
Pubblicazione: (2026)
di: Su, Fei, et al.
Pubblicazione: (2026)
Listening and Seeing Again: Generative Error Correction for Audio-Visual Speech Recognition
di: Liu, Rui, et al.
Pubblicazione: (2025)
di: Liu, Rui, et al.
Pubblicazione: (2025)
Quantitative Analysis of Audio-Visual Tasks: An Information-Theoretic Perspective
di: Chen, Chen, et al.
Pubblicazione: (2024)
di: Chen, Chen, et al.
Pubblicazione: (2024)
AVE Speech: A Comprehensive Multi-Modal Dataset for Speech Recognition Integrating Audio, Visual, and Electromyographic Signals
di: Zhou, Dongliang, et al.
Pubblicazione: (2025)
di: Zhou, Dongliang, et al.
Pubblicazione: (2025)
Efficient Speech Watermarking for Speech Synthesis via Progressive Knowledge Distillation
di: Cui, Yang, et al.
Pubblicazione: (2025)
di: Cui, Yang, et al.
Pubblicazione: (2025)
Zero-Shot End-to-End Spoken Language Understanding via Cross-Modal Selective Self-Training
di: He, Jianfeng, et al.
Pubblicazione: (2023)
di: He, Jianfeng, et al.
Pubblicazione: (2023)
Stimulus Modality Matters: Impact of Perceptual Evaluations from Different Modalities on Speech Emotion Recognition System Performance
di: Chou, Huang-Cheng, et al.
Pubblicazione: (2024)
di: Chou, Huang-Cheng, et al.
Pubblicazione: (2024)
Audio-Thinker: Guiding Audio Language Model When and How to Think via Reinforcement Learning
di: Wu, Shu, et al.
Pubblicazione: (2025)
di: Wu, Shu, et al.
Pubblicazione: (2025)
LaunchpadGPT: Language Model as Music Visualization Designer on Launchpad
di: Xu, Siting, et al.
Pubblicazione: (2023)
di: Xu, Siting, et al.
Pubblicazione: (2023)
AudioSetMix: Enhancing Audio-Language Datasets with LLM-Assisted Augmentations
di: Xu, David
Pubblicazione: (2024)
di: Xu, David
Pubblicazione: (2024)
Documenti analoghi
-
Revise, Reason, and Recognize: LLM-Based Emotion Recognition via Emotion-Specific Prompts and ASR Error Correction
di: Li, Yuanchao, et al.
Pubblicazione: (2024) -
Crossmodal ASR Error Correction with Discrete Speech Units
di: Li, Yuanchao, et al.
Pubblicazione: (2024) -
MF-AED-AEC: Speech Emotion Recognition by Leveraging Multimodal Fusion, Asr Error Detection, and Asr Error Correction
di: He, Jiajun, et al.
Pubblicazione: (2024) -
Semi-Supervised Cognitive State Classification from Speech with Multi-View Pseudo-Labeling
di: Li, Yuanchao, et al.
Pubblicazione: (2024) -
Addressing Emotion Bias in Music Emotion Recognition and Generation with Frechet Audio Distance
di: Li, Yuanchao, et al.
Pubblicazione: (2024)