EmoQ: Speech Emotion Recognition via Speech-Aware Q-Former and Large Language Model
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Yang, Yiqing, Mak, Man-Wai |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
EmoFormer: A Text-Independent Speech Emotion Recognition using a Hybrid Transformer-CNN model
par: Hasan, Rashedul, et autres
Publié: (2025)
par: Hasan, Rashedul, et autres
Publié: (2025)
From Human Speech to Ocean Signals: Transferring Speech Large Models for Underwater Acoustic Target Recognition
par: Huang, Mengcheng, et autres
Publié: (2026)
par: Huang, Mengcheng, et autres
Publié: (2026)
BLSP-Emo: Towards Empathetic Large Speech-Language Models
par: Wang, Chen, et autres
Publié: (2024)
par: Wang, Chen, et autres
Publié: (2024)
VoxEmo: Benchmarking Speech Emotion Recognition with Speech LLMs
par: Zhang, Hezhao, et autres
Publié: (2026)
par: Zhang, Hezhao, et autres
Publié: (2026)
EmoSpeech: A Corpus of Emotionally Rich and Contextually Detailed Speech Annotations
par: Bian, Weizhen, et autres
Publié: (2024)
par: Bian, Weizhen, et autres
Publié: (2024)
Speech Emotion Recognition with ASR Integration
par: Li, Yuanchao
Publié: (2026)
par: Li, Yuanchao
Publié: (2026)
EmoTech: A Multi-modal Speech Emotion Recognition Using Multi-source Low-level Information with Hybrid Recurrent Network
par: Avro, Shamin Bin Habib, et autres
Publié: (2025)
par: Avro, Shamin Bin Habib, et autres
Publié: (2025)
Dataset-Distillation Generative Model for Speech Emotion Recognition
par: Ritter-Gutierrez, Fabian, et autres
Publié: (2024)
par: Ritter-Gutierrez, Fabian, et autres
Publié: (2024)
Emotion-Aware Contrastive Adaptation Network for Source-Free Cross-Corpus Speech Emotion Recognition
par: Zhao, Yan, et autres
Publié: (2024)
par: Zhao, Yan, et autres
Publié: (2024)
Zero-Shot Recognition of Dysarthric Speech Using Commercial Automatic Speech Recognition and Multimodal Large Language Models
par: Alsayegh, Ali, et autres
Publié: (2025)
par: Alsayegh, Ali, et autres
Publié: (2025)
PCQ: Emotion Recognition in Speech via Progressive Channel Querying
par: Wang, Xincheng, et autres
Publié: (2024)
par: Wang, Xincheng, et autres
Publié: (2024)
Attention-weighted Centered Kernel Alignment for Knowledge Distillation in Large Audio-Language Models Applied to Speech Emotion Recognition
par: Yang, Qingran, et autres
Publié: (2026)
par: Yang, Qingran, et autres
Publié: (2026)
Emo-DPO: Controllable Emotional Speech Synthesis through Direct Preference Optimization
par: Gao, Xiaoxue, et autres
Publié: (2024)
par: Gao, Xiaoxue, et autres
Publié: (2024)
EmoSphere-SER: Enhancing Speech Emotion Recognition Through Spherical Representation with Auxiliary Classification
par: Cho, Deok-Hyeon, et autres
Publié: (2025)
par: Cho, Deok-Hyeon, et autres
Publié: (2025)
EmoAttack: Utilizing Emotional Voice Conversion for Speech Backdoor Attacks on Deep Speech Classification Models
par: Yao, Wenhan, et autres
Publié: (2024)
par: Yao, Wenhan, et autres
Publié: (2024)
Emotion Neural Transducer for Fine-Grained Speech Emotion Recognition
par: Shen, Siyuan, et autres
Publié: (2024)
par: Shen, Siyuan, et autres
Publié: (2024)
Testing Correctness, Fairness, and Robustness of Speech Emotion Recognition Models
par: Derington, Anna, et autres
Publié: (2023)
par: Derington, Anna, et autres
Publié: (2023)
EmoSphere-TTS: Emotional Style and Intensity Modeling via Spherical Emotion Vector for Controllable Emotional Text-to-Speech
par: Cho, Deok-Hyeon, et autres
Publié: (2024)
par: Cho, Deok-Hyeon, et autres
Publié: (2024)
Machine Unlearning in Speech Emotion Recognition via Forget Set Alone
par: Ren, Zhao, et autres
Publié: (2025)
par: Ren, Zhao, et autres
Publié: (2025)
DualSpeechLM: Towards Unified Speech Understanding and Generation via Dual Speech Token Modeling with Large Language Models
par: Wang, Yuanyuan, et autres
Publié: (2025)
par: Wang, Yuanyuan, et autres
Publié: (2025)
AmbER$^2$: Dual Ambiguity-Aware Emotion Recognition Applied to Speech and Text
par: Wu, Jingyao, et autres
Publié: (2026)
par: Wu, Jingyao, et autres
Publié: (2026)
NAST: Noise Aware Speech Tokenization for Speech Language Models
par: Messica, Shoval, et autres
Publié: (2024)
par: Messica, Shoval, et autres
Publié: (2024)
THAI Speech Emotion Recognition (THAI-SER) corpus
par: Wongpithayadisai, Jilamika, et autres
Publié: (2025)
par: Wongpithayadisai, Jilamika, et autres
Publié: (2025)
EMO-SUPERB: An In-depth Look at Speech Emotion Recognition
par: Wu, Haibin, et autres
Publié: (2024)
par: Wu, Haibin, et autres
Publié: (2024)
Iterative Prototype Refinement for Ambiguous Speech Emotion Recognition
par: Sun, Haoqin, et autres
Publié: (2024)
par: Sun, Haoqin, et autres
Publié: (2024)
EmoSphere++: Emotion-Controllable Zero-Shot Text-to-Speech via Emotion-Adaptive Spherical Vector
par: Cho, Deok-Hyeon, et autres
Publié: (2024)
par: Cho, Deok-Hyeon, et autres
Publié: (2024)
EMO-RL: Emotion-Rule-Based Reinforcement Learning Enhanced Audio-Language Model for Generalized Speech Emotion Recognition
par: Li, Pengcheng, et autres
Publié: (2025)
par: Li, Pengcheng, et autres
Publié: (2025)
USM-Lite: Quantization and Sparsity Aware Fine-tuning for Speech Recognition with Universal Speech Models
par: Ding, Shaojin, et autres
Publié: (2023)
par: Ding, Shaojin, et autres
Publié: (2023)
Enhancing Speech Large Language Models with Prompt-Aware Mixture of Audio Encoders
par: Shan, Weiqiao, et autres
Publié: (2025)
par: Shan, Weiqiao, et autres
Publié: (2025)
VoxGenesis: Unsupervised Discovery of Latent Speaker Manifold for Speech Synthesis
par: Lin, Weiwei, et autres
Publié: (2024)
par: Lin, Weiwei, et autres
Publié: (2024)
Improving Speech Emotion Recognition in Under-Resourced Languages via Speech-to-Speech Translation with Bootstrapping Data Selection
par: Lin, Hsi-Che, et autres
Publié: (2024)
par: Lin, Hsi-Che, et autres
Publié: (2024)
Streaming Speech Recognition with Decoder-Only Large Language Models and Latency Optimization
par: Wan, Genshun, et autres
Publié: (2026)
par: Wan, Genshun, et autres
Publié: (2026)
SELM: Enhancing Speech Emotion Recognition for Out-of-Domain Scenarios
par: Bukhari, Hazim, et autres
Publié: (2024)
par: Bukhari, Hazim, et autres
Publié: (2024)
S2ST-Omni: Hierarchical Language-Aware SpeechLLM Adaptation for Multilingual Speech-to-Speech Translation
par: Pan, Yu, et autres
Publié: (2025)
par: Pan, Yu, et autres
Publié: (2025)
PEFT-SER: On the Use of Parameter Efficient Transfer Learning Approaches For Speech Emotion Recognition Using Pre-trained Speech Models
par: Feng, Tiantian, et autres
Publié: (2023)
par: Feng, Tiantian, et autres
Publié: (2023)
EmoSteer-TTS: Fine-Grained and Training-Free Emotion-Controllable Text-to-Speech via Activation Steering
par: Xie, Tianxin, et autres
Publié: (2025)
par: Xie, Tianxin, et autres
Publié: (2025)
Multi-Teacher Language-Aware Knowledge Distillation for Multilingual Speech Emotion Recognition
par: Bijoy, Mehedi Hasan, et autres
Publié: (2025)
par: Bijoy, Mehedi Hasan, et autres
Publié: (2025)
Unifying Listener Scoring Scales: Comparison Learning Framework for Speech Quality Assessment and Continuous Speech Emotion Recognition
par: Hu, Cheng-Hung, et autres
Publié: (2025)
par: Hu, Cheng-Hung, et autres
Publié: (2025)
Exploring Local Interpretable Model-Agnostic Explanations for Speech Emotion Recognition with Distribution-Shift
par: Hjuler, Maja J., et autres
Publié: (2025)
par: Hjuler, Maja J., et autres
Publié: (2025)
Can Large Language Models Aid in Annotating Speech Emotional Data? Uncovering New Frontiers
par: Latif, Siddique, et autres
Publié: (2023)
par: Latif, Siddique, et autres
Publié: (2023)
Documents similaires
-
EmoFormer: A Text-Independent Speech Emotion Recognition using a Hybrid Transformer-CNN model
par: Hasan, Rashedul, et autres
Publié: (2025) -
From Human Speech to Ocean Signals: Transferring Speech Large Models for Underwater Acoustic Target Recognition
par: Huang, Mengcheng, et autres
Publié: (2026) -
BLSP-Emo: Towards Empathetic Large Speech-Language Models
par: Wang, Chen, et autres
Publié: (2024) -
VoxEmo: Benchmarking Speech Emotion Recognition with Speech LLMs
par: Zhang, Hezhao, et autres
Publié: (2026) -
EmoSpeech: A Corpus of Emotionally Rich and Contextually Detailed Speech Annotations
par: Bian, Weizhen, et autres
Publié: (2024)