Stimulus Modality Matters: Impact of Perceptual Evaluations from Different Modalities on Speech Emotion Recognition System Performance
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chou, Huang-Cheng, Wu, Haibin, Lee, Hung-yi, Lee, Chi-Chun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
FoVNet: Configurable Field-of-View Speech Enhancement with Low Computation and Distortion for Smart Glasses
von: Xu, Zhongweiyang, et al.
Veröffentlicht: (2024)
von: Xu, Zhongweiyang, et al.
Veröffentlicht: (2024)
Robust LLM-based Audio-Visual Speech Recognition with Sparse Modality Alignment and Visual Unit-Guided Refinement
von: Su, Fei, et al.
Veröffentlicht: (2026)
von: Su, Fei, et al.
Veröffentlicht: (2026)
Voice Evaluation of Reasoning Ability: Diagnosing the Modality-Induced Performance Gap
von: Lin, Yueqian, et al.
Veröffentlicht: (2025)
von: Lin, Yueqian, et al.
Veröffentlicht: (2025)
Singing Voice Graph Modeling for SingFake Detection
von: Chen, Xuanjun, et al.
Veröffentlicht: (2024)
von: Chen, Xuanjun, et al.
Veröffentlicht: (2024)
EMO-SUPERB: An In-depth Look at Speech Emotion Recognition
von: Wu, Haibin, et al.
Veröffentlicht: (2024)
von: Wu, Haibin, et al.
Veröffentlicht: (2024)
Recent Advances in Discrete Speech Tokens: A Review
von: Guo, Yiwei, et al.
Veröffentlicht: (2025)
von: Guo, Yiwei, et al.
Veröffentlicht: (2025)
VoxEmo: Benchmarking Speech Emotion Recognition with Speech LLMs
von: Zhang, Hezhao, et al.
Veröffentlicht: (2026)
von: Zhang, Hezhao, et al.
Veröffentlicht: (2026)
ArrayDPS: Unsupervised Blind Speech Separation with a Diffusion Prior
von: Xu, Zhongweiyang, et al.
Veröffentlicht: (2025)
von: Xu, Zhongweiyang, et al.
Veröffentlicht: (2025)
EMO-Codec: An In-Depth Look at Emotion Preservation capacity of Legacy and Neural Codec Models With Subjective and Objective Evaluations
von: Ren, Wenze, et al.
Veröffentlicht: (2024)
von: Ren, Wenze, et al.
Veröffentlicht: (2024)
Robust Wake Word Spotting With Frame-Level Cross-Modal Attention Based Audio-Visual Conformer
von: Wang, Haoxu, et al.
Veröffentlicht: (2024)
von: Wang, Haoxu, et al.
Veröffentlicht: (2024)
A Survey on Multimodal Music Emotion Recognition
von: Liyanarachchi, Rashini, et al.
Veröffentlicht: (2025)
von: Liyanarachchi, Rashini, et al.
Veröffentlicht: (2025)
A Unified Framework for Modality-Agnostic Deepfakes Detection
von: Yu, Cai, et al.
Veröffentlicht: (2023)
von: Yu, Cai, et al.
Veröffentlicht: (2023)
VERSA: A Versatile Evaluation Toolkit for Speech, Audio, and Music
von: Shi, Jiatong, et al.
Veröffentlicht: (2024)
von: Shi, Jiatong, et al.
Veröffentlicht: (2024)
Sync-TVA: A Graph-Attention Framework for Multimodal Emotion Recognition with Cross-Modal Fusion
von: Deng, Zeyu, et al.
Veröffentlicht: (2025)
von: Deng, Zeyu, et al.
Veröffentlicht: (2025)
Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities
von: Sudarsanam, Parthasaarathy, et al.
Veröffentlicht: (2025)
von: Sudarsanam, Parthasaarathy, et al.
Veröffentlicht: (2025)
Bridging The Multi-Modality Gaps of Audio, Visual and Linguistic for Speech Enhancement
von: Lin, Meng-Ping, et al.
Veröffentlicht: (2025)
von: Lin, Meng-Ping, et al.
Veröffentlicht: (2025)
Multimodal Emotion Recognition from Raw Audio with Sinc-convolution
von: Zhang, Xiaohui, et al.
Veröffentlicht: (2024)
von: Zhang, Xiaohui, et al.
Veröffentlicht: (2024)
A Survey on Cross-Modal Interaction Between Music and Multimodal Data
von: Li, Sifei, et al.
Veröffentlicht: (2025)
von: Li, Sifei, et al.
Veröffentlicht: (2025)
Emo-bias: A Large Scale Evaluation of Social Bias on Speech Emotion Recognition
von: Lin, Yi-Cheng, et al.
Veröffentlicht: (2024)
von: Lin, Yi-Cheng, et al.
Veröffentlicht: (2024)
Mitigating Subgroup Disparities in Multi-Label Speech Emotion Recognition: A Pseudo-Labeling and Unsupervised Learning Approach
von: Lin, Yi-Cheng, et al.
Veröffentlicht: (2025)
von: Lin, Yi-Cheng, et al.
Veröffentlicht: (2025)
Joint Fullband-Subband Modeling for High-Resolution SingFake Detection
von: Chen, Xuanjun, et al.
Veröffentlicht: (2026)
von: Chen, Xuanjun, et al.
Veröffentlicht: (2026)
Cross-Modal Watermarking for Authentic Audio Recovery and Tamper Localization in Synthesized Audiovisual Forgeries
von: Kim, Minyoung, et al.
Veröffentlicht: (2025)
von: Kim, Minyoung, et al.
Veröffentlicht: (2025)
LCB-net: Long-Context Biasing for Audio-Visual Speech Recognition
von: Yu, Fan, et al.
Veröffentlicht: (2024)
von: Yu, Fan, et al.
Veröffentlicht: (2024)
Human-Inspired Computing for Robust and Efficient Audio-Visual Speech Recognition
von: Liu, Qianhui, et al.
Veröffentlicht: (2024)
von: Liu, Qianhui, et al.
Veröffentlicht: (2024)
Improving Speech Emotion Recognition in Under-Resourced Languages via Speech-to-Speech Translation with Bootstrapping Data Selection
von: Lin, Hsi-Che, et al.
Veröffentlicht: (2024)
von: Lin, Hsi-Che, et al.
Veröffentlicht: (2024)
MSAC: Multiple Speech Attribute Control Method for Reliable Speech Emotion Recognition
von: Pan, Yu, et al.
Veröffentlicht: (2023)
von: Pan, Yu, et al.
Veröffentlicht: (2023)
Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation
von: Rahimi, Akam, et al.
Veröffentlicht: (2025)
von: Rahimi, Akam, et al.
Veröffentlicht: (2025)
Analyzable Chain-of-Musical-Thought Prompting for High-Fidelity Music Generation
von: Lam, Max W. Y., et al.
Veröffentlicht: (2025)
von: Lam, Max W. Y., et al.
Veröffentlicht: (2025)
ARECHO: Autoregressive Evaluation via Chain-Based Hypothesis Optimization for Speech Multi-Metric Estimation
von: Shi, Jiatong, et al.
Veröffentlicht: (2025)
von: Shi, Jiatong, et al.
Veröffentlicht: (2025)
CodecFake: Enhancing Anti-Spoofing Models Against Deepfake Audios from Codec-Based Speech Synthesis Systems
von: Wu, Haibin, et al.
Veröffentlicht: (2024)
von: Wu, Haibin, et al.
Veröffentlicht: (2024)
AVE Speech: A Comprehensive Multi-Modal Dataset for Speech Recognition Integrating Audio, Visual, and Electromyographic Signals
von: Zhou, Dongliang, et al.
Veröffentlicht: (2025)
von: Zhou, Dongliang, et al.
Veröffentlicht: (2025)
Continuous Speech Tokens Makes LLMs Robust Multi-Modality Learners
von: Yuan, Ze, et al.
Veröffentlicht: (2024)
von: Yuan, Ze, et al.
Veröffentlicht: (2024)
WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling
von: Ji, Shengpeng, et al.
Veröffentlicht: (2024)
von: Ji, Shengpeng, et al.
Veröffentlicht: (2024)
Early Joint Learning of Emotion Information Makes MultiModal Model Understand You Better
von: Ge, Mengying, et al.
Veröffentlicht: (2024)
von: Ge, Mengying, et al.
Veröffentlicht: (2024)
Low-latency Speech Enhancement via Speech Token Generation
von: Xue, Huaying, et al.
Veröffentlicht: (2023)
von: Xue, Huaying, et al.
Veröffentlicht: (2023)
Learning Temporal Resolution in Spectrogram for Audio Classification
von: Liu, Haohe, et al.
Veröffentlicht: (2022)
von: Liu, Haohe, et al.
Veröffentlicht: (2022)
AudioLDM 2: Learning Holistic Audio Generation with Self-supervised Pretraining
von: Liu, Haohe, et al.
Veröffentlicht: (2023)
von: Liu, Haohe, et al.
Veröffentlicht: (2023)
SemantiCodec: An Ultra Low Bitrate Semantic Audio Codec for General Sound
von: Liu, Haohe, et al.
Veröffentlicht: (2024)
von: Liu, Haohe, et al.
Veröffentlicht: (2024)
SynthTab: Leveraging Synthesized Data for Guitar Tablature Transcription
von: Zang, Yongyi, et al.
Veröffentlicht: (2023)
von: Zang, Yongyi, et al.
Veröffentlicht: (2023)
Episodic fine-tuning prototypical networks for optimization-based few-shot learning: Application to audio classification
von: Zhuang, Xuanyu, et al.
Veröffentlicht: (2024)
von: Zhuang, Xuanyu, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
FoVNet: Configurable Field-of-View Speech Enhancement with Low Computation and Distortion for Smart Glasses
von: Xu, Zhongweiyang, et al.
Veröffentlicht: (2024) -
Robust LLM-based Audio-Visual Speech Recognition with Sparse Modality Alignment and Visual Unit-Guided Refinement
von: Su, Fei, et al.
Veröffentlicht: (2026) -
Voice Evaluation of Reasoning Ability: Diagnosing the Modality-Induced Performance Gap
von: Lin, Yueqian, et al.
Veröffentlicht: (2025) -
Singing Voice Graph Modeling for SingFake Detection
von: Chen, Xuanjun, et al.
Veröffentlicht: (2024) -
EMO-SUPERB: An In-depth Look at Speech Emotion Recognition
von: Wu, Haibin, et al.
Veröffentlicht: (2024)