Gen-SER: When the generative model meets speech emotion recognition
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Wang, Taihui, Zhao, Jinzheng, Chen, Rilin, Lei, Tong, Wang, Wenwu, Yu, Dong |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Target matching based generative model for speech enhancement
par: Wang, Taihui, et autres
Publié: (2025)
par: Wang, Taihui, et autres
Publié: (2025)
AudioRAG+: Feedback-driven Retrieval-augmented Audio Generation with Large Audio Language Models
par: Zhao, Junqi, et autres
Publié: (2025)
par: Zhao, Junqi, et autres
Publié: (2025)
Heterogeneous bimodal attention fusion for speech emotion recognition
par: Luo, Jiachen, et autres
Publié: (2025)
par: Luo, Jiachen, et autres
Publié: (2025)
Graph-based multi-Feature fusion method for speech emotion recognition
par: Liu, Xueyu, et autres
Publié: (2024)
par: Liu, Xueyu, et autres
Publié: (2024)
Beyond saliency: enhancing explanation of speech emotion recognition with expert-referenced acoustic cues
par: Nasr, Seham, et autres
Publié: (2025)
par: Nasr, Seham, et autres
Publié: (2025)
Scalable Neural Vocoder from Range-Null Space Decomposition
par: Li, Andong, et autres
Publié: (2026)
par: Li, Andong, et autres
Publié: (2026)
BridgeVoC: Revitalizing Neural Vocoder from a Restoration Perspective
par: Li, Andong, et autres
Publié: (2025)
par: Li, Andong, et autres
Publié: (2025)
A vector quantized masked autoencoder for audiovisual speech emotion recognition
par: Sadok, Samir, et autres
Publié: (2023)
par: Sadok, Samir, et autres
Publié: (2023)
Charting 15 years of progress in deep learning for speech emotion recognition: A replication study
par: Triantafyllopoulos, Andreas, et autres
Publié: (2025)
par: Triantafyllopoulos, Andreas, et autres
Publié: (2025)
Enhancing CTC-based speech recognition with diverse modeling units
par: Han, Shiyi, et autres
Publié: (2024)
par: Han, Shiyi, et autres
Publié: (2024)
Fusion approaches for emotion recognition from speech using acoustic and text-based features
par: Pepino, Leonardo, et autres
Publié: (2024)
par: Pepino, Leonardo, et autres
Publié: (2024)
learning discriminative features from spectrograms using center loss for speech emotion recognition
par: Dai, Dongyang, et autres
Publié: (2025)
par: Dai, Dongyang, et autres
Publié: (2025)
Robust fine-tuning of speech recognition models via model merging: application to disordered speech
par: Ducorroy, Alexandre, et autres
Publié: (2025)
par: Ducorroy, Alexandre, et autres
Publié: (2025)
Explainable speech emotion recognition through attentive pooling: insights from attention-based temporal localization
par: Leygue, Tahitoa, et autres
Publié: (2025)
par: Leygue, Tahitoa, et autres
Publié: (2025)
Learning Neural Vocoder from Range-Null Space Decomposition
par: Li, Andong, et autres
Publié: (2025)
par: Li, Andong, et autres
Publié: (2025)
Advancing LLM-based phoneme-to-grapheme for multilingual speech recognition
par: Dong, Lukuang, et autres
Publié: (2026)
par: Dong, Lukuang, et autres
Publié: (2026)
Video-to-Audio Generation with Fine-grained Temporal Semantics
par: Hu, Yuchen, et autres
Publié: (2024)
par: Hu, Yuchen, et autres
Publié: (2024)
Fish Tracking, Counting, and Behaviour Analysis in Digital Aquaculture: A Comprehensive Survey
par: Cui, Meng, et autres
Publié: (2024)
par: Cui, Meng, et autres
Publié: (2024)
TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet
par: Jeong, Jaeseok, et autres
Publié: (2025)
par: Jeong, Jaeseok, et autres
Publié: (2025)
AS-70: A Mandarin stuttered speech dataset for automatic speech recognition and stuttering event detection
par: Gong, Rong, et autres
Publié: (2024)
par: Gong, Rong, et autres
Publié: (2024)
Introduction to speech recognition
par: Dauphin, Gabriel
Publié: (2024)
par: Dauphin, Gabriel
Publié: (2024)
Multi-channel multi-speaker transformer for speech recognition
par: Yifan, Guo, et autres
Publié: (2026)
par: Yifan, Guo, et autres
Publié: (2026)
Improving child speech recognition with augmented child-like speech
par: Zhang, Yuanyuan, et autres
Publié: (2024)
par: Zhang, Yuanyuan, et autres
Publié: (2024)
Language model integration based on memory control for sequence to sequence speech recognition
par: Cho, Jaejin, et autres
Publié: (2018)
par: Cho, Jaejin, et autres
Publié: (2018)
Phoneme-based speech recognition driven by large language models and sampling marginalization
par: Ma, Te, et autres
Publié: (2025)
par: Ma, Te, et autres
Publié: (2025)
AudioTurbo: Fast Text-to-Audio Generation with Rectified Diffusion
par: Zhao, Junqi, et autres
Publié: (2025)
par: Zhao, Junqi, et autres
Publié: (2025)
From Continuous to Discrete: Cross-Domain Collaborative General Speech Enhancement via Hierarchical Language Models
par: Mu, Zhaoxi, et autres
Publié: (2025)
par: Mu, Zhaoxi, et autres
Publié: (2025)
SSR-Speech: Towards Stable, Safe and Robust Zero-shot Text-based Speech Editing and Synthesis
par: Wang, Helin, et autres
Publié: (2024)
par: Wang, Helin, et autres
Publié: (2024)
Voxtlm: unified decoder-only models for consolidating speech recognition/synthesis and speech/text continuation tasks
par: Maiti, Soumi, et autres
Publié: (2023)
par: Maiti, Soumi, et autres
Publié: (2023)
Universal Sound Separation with Self-Supervised Audio Masked Autoencoder
par: Zhao, Junqi, et autres
Publié: (2024)
par: Zhao, Junqi, et autres
Publié: (2024)
Training chord recognition models on artificially generated audio
par: Majchrzak, Martyna, et autres
Publié: (2025)
par: Majchrzak, Martyna, et autres
Publié: (2025)
Index-MSR: A high-efficiency multimodal fusion framework for speech recognition
par: Chen, Jinming, et autres
Publié: (2025)
par: Chen, Jinming, et autres
Publié: (2025)
Keyword spotting using convolutional neural network for speech recognition in Hindi
par: Bharti, Saru, et autres
Publié: (2026)
par: Bharti, Saru, et autres
Publié: (2026)
Exploring the limits of decoder-only models trained on public speech recognition corpora
par: Gupta, Ankit, et autres
Publié: (2024)
par: Gupta, Ankit, et autres
Publié: (2024)
A unified multichannel far-field speech recognition system: combining neural beamforming with attention based end-to-end model
par: Zhao, Dongdi, et autres
Publié: (2024)
par: Zhao, Dongdi, et autres
Publié: (2024)
StemGen: A music generation model that listens
par: Parker, Julian D., et autres
Publié: (2023)
par: Parker, Julian D., et autres
Publié: (2023)
Region-Specific Audio Tagging for Spatial Sound
par: Zhao, Jinzheng, et autres
Publié: (2025)
par: Zhao, Jinzheng, et autres
Publié: (2025)
SMRU: Split-and-Merge Recurrent-based UNet for Acoustic Echo Cancellation and Noise Suppression
par: Sun, Zhihang, et autres
Publié: (2024)
par: Sun, Zhihang, et autres
Publié: (2024)
STA-V2A: Video-to-Audio Generation with Semantic and Temporal Alignment
par: Ren, Yong, et autres
Publié: (2024)
par: Ren, Yong, et autres
Publié: (2024)
Towards interpretable emotion recognition: Identifying key features with machine learning
par: Kaloga, Yacouba, et autres
Publié: (2025)
par: Kaloga, Yacouba, et autres
Publié: (2025)
Documents similaires
-
Target matching based generative model for speech enhancement
par: Wang, Taihui, et autres
Publié: (2025) -
AudioRAG+: Feedback-driven Retrieval-augmented Audio Generation with Large Audio Language Models
par: Zhao, Junqi, et autres
Publié: (2025) -
Heterogeneous bimodal attention fusion for speech emotion recognition
par: Luo, Jiachen, et autres
Publié: (2025) -
Graph-based multi-Feature fusion method for speech emotion recognition
par: Liu, Xueyu, et autres
Publié: (2024) -
Beyond saliency: enhancing explanation of speech emotion recognition with expert-referenced acoustic cues
par: Nasr, Seham, et autres
Publié: (2025)