Scaling Ambiguity: Augmenting Human Annotation in Speech Emotion Recognition with Audio-Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Wenda, Jin, Hongyu, Wang, Siyi, Wei, Zhiqiang, Dang, Ting |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Decoding Ambiguous Emotions with Test-Time Scaling in Audio-Language Models
von: Jia, Hong, et al.
Veröffentlicht: (2026)
von: Jia, Hong, et al.
Veröffentlicht: (2026)
Scaling Auditory Cognition via Test-Time Compute in Audio Language Models
von: Dang, Ting, et al.
Veröffentlicht: (2025)
von: Dang, Ting, et al.
Veröffentlicht: (2025)
Disentangling Reasoning in Large Audio-Language Models for Ambiguous Emotion Prediction
von: Yu, Xiaofeng, et al.
Veröffentlicht: (2026)
von: Yu, Xiaofeng, et al.
Veröffentlicht: (2026)
Token-Level Logits Matter: A Closer Look at Speech Foundation Models for Ambiguous Emotion Recognition
von: Halim, Jule Valendo, et al.
Veröffentlicht: (2025)
von: Halim, Jule Valendo, et al.
Veröffentlicht: (2025)
Listening and Seeing Again: Generative Error Correction for Audio-Visual Speech Recognition
von: Liu, Rui, et al.
Veröffentlicht: (2025)
von: Liu, Rui, et al.
Veröffentlicht: (2025)
Color-based Emotion Representation for Speech Emotion Recognition
von: Nagase, Ryotaro, et al.
Veröffentlicht: (2026)
von: Nagase, Ryotaro, et al.
Veröffentlicht: (2026)
EmoSpeech: A Corpus of Emotionally Rich and Contextually Detailed Speech Annotations
von: Bian, Weizhen, et al.
Veröffentlicht: (2024)
von: Bian, Weizhen, et al.
Veröffentlicht: (2024)
Persian Speech Emotion Recognition by Fine-Tuning Transformers
von: Shayaninasab, Minoo, et al.
Veröffentlicht: (2024)
von: Shayaninasab, Minoo, et al.
Veröffentlicht: (2024)
Enhancing Speech Emotion Recognition through Segmental Average Pooling of Self-Supervised Learning Features
von: Hyeon, Jonghwan, et al.
Veröffentlicht: (2024)
von: Hyeon, Jonghwan, et al.
Veröffentlicht: (2024)
Audio Codec Augmentation for Robust Collaborative Watermarking of Speech Synthesis
von: Juvela, Lauri, et al.
Veröffentlicht: (2024)
von: Juvela, Lauri, et al.
Veröffentlicht: (2024)
Rethinking Continual Learning for Speech and Audio: A Representation-Centric Taxonomy and Open Problems
von: Xiao, Yang, et al.
Veröffentlicht: (2026)
von: Xiao, Yang, et al.
Veröffentlicht: (2026)
Efficient Finetuning for Dimensional Speech Emotion Recognition in the Age of Transformers
von: Sampath, Aneesha, et al.
Veröffentlicht: (2025)
von: Sampath, Aneesha, et al.
Veröffentlicht: (2025)
Improvement and Implementation of a Speech Emotion Recognition Model Based on Dual-Layer LSTM
von: Yang, Xiaoran, et al.
Veröffentlicht: (2024)
von: Yang, Xiaoran, et al.
Veröffentlicht: (2024)
Speech Emotion Recognition Using MFCC Features and LSTM-Based Deep Learning Model
von: Oluwademilade, Adelekun, et al.
Veröffentlicht: (2026)
von: Oluwademilade, Adelekun, et al.
Veröffentlicht: (2026)
Breaking Resource Barriers in Speech Emotion Recognition via Data Distillation
von: Chang, Yi, et al.
Veröffentlicht: (2024)
von: Chang, Yi, et al.
Veröffentlicht: (2024)
Learning Physiology-Informed Vocal Spectrotemporal Representations for Speech Emotion Recognition
von: Zhang, Xu, et al.
Veröffentlicht: (2026)
von: Zhang, Xu, et al.
Veröffentlicht: (2026)
Toward Efficient Speech Emotion Recognition via Spectral Learning and Attention
von: Lee, HyeYoung, et al.
Veröffentlicht: (2025)
von: Lee, HyeYoung, et al.
Veröffentlicht: (2025)
ABHINAYA -- A System for Speech Emotion Recognition In Naturalistic Conditions Challenge
von: Dutta, Soumya, et al.
Veröffentlicht: (2025)
von: Dutta, Soumya, et al.
Veröffentlicht: (2025)
Active Learning with Task Adaptation Pre-training for Speech Emotion Recognition
von: Li, Dongyuan, et al.
Veröffentlicht: (2024)
von: Li, Dongyuan, et al.
Veröffentlicht: (2024)
Multi-Loss Learning for Speech Emotion Recognition with Energy-Adaptive Mixup and Frame-Level Attention
von: Wang, Cong, et al.
Veröffentlicht: (2025)
von: Wang, Cong, et al.
Veröffentlicht: (2025)
VoxPrivacy: A Benchmark for Evaluating Interactional Privacy of Speech Language Models
von: Wang, Yuxiang, et al.
Veröffentlicht: (2026)
von: Wang, Yuxiang, et al.
Veröffentlicht: (2026)
Are you sure? Analysing Uncertainty Quantification Approaches for Real-world Speech Emotion Recognition
von: Schrüfer, Oliver, et al.
Veröffentlicht: (2024)
von: Schrüfer, Oliver, et al.
Veröffentlicht: (2024)
Explaining Deep Learning Embeddings for Speech Emotion Recognition by Predicting Interpretable Acoustic Features
von: Dixit, Satvik, et al.
Veröffentlicht: (2024)
von: Dixit, Satvik, et al.
Veröffentlicht: (2024)
MLCA-AVSR: Multi-Layer Cross Attention Fusion based Audio-Visual Speech Recognition
von: Wang, He, et al.
Veröffentlicht: (2024)
von: Wang, He, et al.
Veröffentlicht: (2024)
The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition
von: Gao, Ming, et al.
Veröffentlicht: (2025)
von: Gao, Ming, et al.
Veröffentlicht: (2025)
Exploring Self-Supervised Multi-view Contrastive Learning for Speech Emotion Recognition with Limited Annotations
von: Khaertdinov, Bulat, et al.
Veröffentlicht: (2024)
von: Khaertdinov, Bulat, et al.
Veröffentlicht: (2024)
EmoSphere-SER: Enhancing Speech Emotion Recognition Through Spherical Representation with Auxiliary Classification
von: Cho, Deok-Hyeon, et al.
Veröffentlicht: (2025)
von: Cho, Deok-Hyeon, et al.
Veröffentlicht: (2025)
Cross-Corpus Validation of Speech Emotion Recognition in Urdu using Domain-Knowledge Acoustic Features
von: Talpur, Unzela, et al.
Veröffentlicht: (2025)
von: Talpur, Unzela, et al.
Veröffentlicht: (2025)
MFHCA: Enhancing Speech Emotion Recognition Via Multi-Spatial Fusion and Hierarchical Cooperative Attention
von: Jiao, Xinxin, et al.
Veröffentlicht: (2024)
von: Jiao, Xinxin, et al.
Veröffentlicht: (2024)
MATER: Multi-level Acoustic and Textual Emotion Representation for Interpretable Speech Emotion Recognition
von: Jon, Hyo Jin, et al.
Veröffentlicht: (2025)
von: Jon, Hyo Jin, et al.
Veröffentlicht: (2025)
WavRAG: Audio-Integrated Retrieval Augmented Generation for Spoken Dialogue Models
von: Chen, Yifu, et al.
Veröffentlicht: (2025)
von: Chen, Yifu, et al.
Veröffentlicht: (2025)
Zero-Shot Parkinson's Disease Detection from Speech: Comparing Large Audio and Language Models
von: Kabir, Muhammad Ashad, et al.
Veröffentlicht: (2026)
von: Kabir, Muhammad Ashad, et al.
Veröffentlicht: (2026)
IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System
von: Deng, Wei, et al.
Veröffentlicht: (2025)
von: Deng, Wei, et al.
Veröffentlicht: (2025)
MSAC: Multiple Speech Attribute Control Method for Reliable Speech Emotion Recognition
von: Pan, Yu, et al.
Veröffentlicht: (2023)
von: Pan, Yu, et al.
Veröffentlicht: (2023)
Tiny-Align: Bridging Automatic Speech Recognition and Large Language Model on the Edge
von: Qin, Ruiyang, et al.
Veröffentlicht: (2024)
von: Qin, Ruiyang, et al.
Veröffentlicht: (2024)
Inference-time Scaling for Diffusion-based Audio Super-resolution
von: Jin, Yizhu, et al.
Veröffentlicht: (2025)
von: Jin, Yizhu, et al.
Veröffentlicht: (2025)
Searching for Effective Preprocessing Method and CNN-based Architecture with Efficient Channel Attention on Speech Emotion Recognition
von: Kim, Byunggun, et al.
Veröffentlicht: (2024)
von: Kim, Byunggun, et al.
Veröffentlicht: (2024)
Sommelier: Scalable Open Multi-turn Audio Pre-processing for Full-duplex Speech Language Models
von: Jung, Kyudan, et al.
Veröffentlicht: (2026)
von: Jung, Kyudan, et al.
Veröffentlicht: (2026)
Self-supervised ASR Models and Features For Dysarthric and Elderly Speech Recognition
von: Hu, Shujie, et al.
Veröffentlicht: (2024)
von: Hu, Shujie, et al.
Veröffentlicht: (2024)
Gibberish is All You Need for Membership Inference Detection in Contrastive Language-Audio Pretraining
von: Cheng, Ruoxi, et al.
Veröffentlicht: (2024)
von: Cheng, Ruoxi, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Decoding Ambiguous Emotions with Test-Time Scaling in Audio-Language Models
von: Jia, Hong, et al.
Veröffentlicht: (2026) -
Scaling Auditory Cognition via Test-Time Compute in Audio Language Models
von: Dang, Ting, et al.
Veröffentlicht: (2025) -
Disentangling Reasoning in Large Audio-Language Models for Ambiguous Emotion Prediction
von: Yu, Xiaofeng, et al.
Veröffentlicht: (2026) -
Token-Level Logits Matter: A Closer Look at Speech Foundation Models for Ambiguous Emotion Recognition
von: Halim, Jule Valendo, et al.
Veröffentlicht: (2025) -
Listening and Seeing Again: Generative Error Correction for Audio-Visual Speech Recognition
von: Liu, Rui, et al.
Veröffentlicht: (2025)