Gen-SER: When the generative model meets speech emotion recognition
Fuente:
arXiv
Guardado en:
| Autores principales: | Wang, Taihui, Zhao, Jinzheng, Chen, Rilin, Lei, Tong, Wang, Wenwu, Yu, Dong |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Target matching based generative model for speech enhancement
por: Wang, Taihui, et al.
Publicado: (2025)
por: Wang, Taihui, et al.
Publicado: (2025)
AudioRAG+: Feedback-driven Retrieval-augmented Audio Generation with Large Audio Language Models
por: Zhao, Junqi, et al.
Publicado: (2025)
por: Zhao, Junqi, et al.
Publicado: (2025)
Heterogeneous bimodal attention fusion for speech emotion recognition
por: Luo, Jiachen, et al.
Publicado: (2025)
por: Luo, Jiachen, et al.
Publicado: (2025)
Graph-based multi-Feature fusion method for speech emotion recognition
por: Liu, Xueyu, et al.
Publicado: (2024)
por: Liu, Xueyu, et al.
Publicado: (2024)
Beyond saliency: enhancing explanation of speech emotion recognition with expert-referenced acoustic cues
por: Nasr, Seham, et al.
Publicado: (2025)
por: Nasr, Seham, et al.
Publicado: (2025)
Scalable Neural Vocoder from Range-Null Space Decomposition
por: Li, Andong, et al.
Publicado: (2026)
por: Li, Andong, et al.
Publicado: (2026)
BridgeVoC: Revitalizing Neural Vocoder from a Restoration Perspective
por: Li, Andong, et al.
Publicado: (2025)
por: Li, Andong, et al.
Publicado: (2025)
A vector quantized masked autoencoder for audiovisual speech emotion recognition
por: Sadok, Samir, et al.
Publicado: (2023)
por: Sadok, Samir, et al.
Publicado: (2023)
Charting 15 years of progress in deep learning for speech emotion recognition: A replication study
por: Triantafyllopoulos, Andreas, et al.
Publicado: (2025)
por: Triantafyllopoulos, Andreas, et al.
Publicado: (2025)
Enhancing CTC-based speech recognition with diverse modeling units
por: Han, Shiyi, et al.
Publicado: (2024)
por: Han, Shiyi, et al.
Publicado: (2024)
Fusion approaches for emotion recognition from speech using acoustic and text-based features
por: Pepino, Leonardo, et al.
Publicado: (2024)
por: Pepino, Leonardo, et al.
Publicado: (2024)
learning discriminative features from spectrograms using center loss for speech emotion recognition
por: Dai, Dongyang, et al.
Publicado: (2025)
por: Dai, Dongyang, et al.
Publicado: (2025)
Robust fine-tuning of speech recognition models via model merging: application to disordered speech
por: Ducorroy, Alexandre, et al.
Publicado: (2025)
por: Ducorroy, Alexandre, et al.
Publicado: (2025)
Explainable speech emotion recognition through attentive pooling: insights from attention-based temporal localization
por: Leygue, Tahitoa, et al.
Publicado: (2025)
por: Leygue, Tahitoa, et al.
Publicado: (2025)
Learning Neural Vocoder from Range-Null Space Decomposition
por: Li, Andong, et al.
Publicado: (2025)
por: Li, Andong, et al.
Publicado: (2025)
Advancing LLM-based phoneme-to-grapheme for multilingual speech recognition
por: Dong, Lukuang, et al.
Publicado: (2026)
por: Dong, Lukuang, et al.
Publicado: (2026)
Video-to-Audio Generation with Fine-grained Temporal Semantics
por: Hu, Yuchen, et al.
Publicado: (2024)
por: Hu, Yuchen, et al.
Publicado: (2024)
Fish Tracking, Counting, and Behaviour Analysis in Digital Aquaculture: A Comprehensive Survey
por: Cui, Meng, et al.
Publicado: (2024)
por: Cui, Meng, et al.
Publicado: (2024)
TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet
por: Jeong, Jaeseok, et al.
Publicado: (2025)
por: Jeong, Jaeseok, et al.
Publicado: (2025)
AS-70: A Mandarin stuttered speech dataset for automatic speech recognition and stuttering event detection
por: Gong, Rong, et al.
Publicado: (2024)
por: Gong, Rong, et al.
Publicado: (2024)
Introduction to speech recognition
por: Dauphin, Gabriel
Publicado: (2024)
por: Dauphin, Gabriel
Publicado: (2024)
Multi-channel multi-speaker transformer for speech recognition
por: Yifan, Guo, et al.
Publicado: (2026)
por: Yifan, Guo, et al.
Publicado: (2026)
Improving child speech recognition with augmented child-like speech
por: Zhang, Yuanyuan, et al.
Publicado: (2024)
por: Zhang, Yuanyuan, et al.
Publicado: (2024)
Language model integration based on memory control for sequence to sequence speech recognition
por: Cho, Jaejin, et al.
Publicado: (2018)
por: Cho, Jaejin, et al.
Publicado: (2018)
Phoneme-based speech recognition driven by large language models and sampling marginalization
por: Ma, Te, et al.
Publicado: (2025)
por: Ma, Te, et al.
Publicado: (2025)
AudioTurbo: Fast Text-to-Audio Generation with Rectified Diffusion
por: Zhao, Junqi, et al.
Publicado: (2025)
por: Zhao, Junqi, et al.
Publicado: (2025)
From Continuous to Discrete: Cross-Domain Collaborative General Speech Enhancement via Hierarchical Language Models
por: Mu, Zhaoxi, et al.
Publicado: (2025)
por: Mu, Zhaoxi, et al.
Publicado: (2025)
SSR-Speech: Towards Stable, Safe and Robust Zero-shot Text-based Speech Editing and Synthesis
por: Wang, Helin, et al.
Publicado: (2024)
por: Wang, Helin, et al.
Publicado: (2024)
Voxtlm: unified decoder-only models for consolidating speech recognition/synthesis and speech/text continuation tasks
por: Maiti, Soumi, et al.
Publicado: (2023)
por: Maiti, Soumi, et al.
Publicado: (2023)
Universal Sound Separation with Self-Supervised Audio Masked Autoencoder
por: Zhao, Junqi, et al.
Publicado: (2024)
por: Zhao, Junqi, et al.
Publicado: (2024)
Training chord recognition models on artificially generated audio
por: Majchrzak, Martyna, et al.
Publicado: (2025)
por: Majchrzak, Martyna, et al.
Publicado: (2025)
Index-MSR: A high-efficiency multimodal fusion framework for speech recognition
por: Chen, Jinming, et al.
Publicado: (2025)
por: Chen, Jinming, et al.
Publicado: (2025)
Keyword spotting using convolutional neural network for speech recognition in Hindi
por: Bharti, Saru, et al.
Publicado: (2026)
por: Bharti, Saru, et al.
Publicado: (2026)
Exploring the limits of decoder-only models trained on public speech recognition corpora
por: Gupta, Ankit, et al.
Publicado: (2024)
por: Gupta, Ankit, et al.
Publicado: (2024)
A unified multichannel far-field speech recognition system: combining neural beamforming with attention based end-to-end model
por: Zhao, Dongdi, et al.
Publicado: (2024)
por: Zhao, Dongdi, et al.
Publicado: (2024)
StemGen: A music generation model that listens
por: Parker, Julian D., et al.
Publicado: (2023)
por: Parker, Julian D., et al.
Publicado: (2023)
Region-Specific Audio Tagging for Spatial Sound
por: Zhao, Jinzheng, et al.
Publicado: (2025)
por: Zhao, Jinzheng, et al.
Publicado: (2025)
SMRU: Split-and-Merge Recurrent-based UNet for Acoustic Echo Cancellation and Noise Suppression
por: Sun, Zhihang, et al.
Publicado: (2024)
por: Sun, Zhihang, et al.
Publicado: (2024)
STA-V2A: Video-to-Audio Generation with Semantic and Temporal Alignment
por: Ren, Yong, et al.
Publicado: (2024)
por: Ren, Yong, et al.
Publicado: (2024)
Towards interpretable emotion recognition: Identifying key features with machine learning
por: Kaloga, Yacouba, et al.
Publicado: (2025)
por: Kaloga, Yacouba, et al.
Publicado: (2025)
Ejemplares similares
-
Target matching based generative model for speech enhancement
por: Wang, Taihui, et al.
Publicado: (2025) -
AudioRAG+: Feedback-driven Retrieval-augmented Audio Generation with Large Audio Language Models
por: Zhao, Junqi, et al.
Publicado: (2025) -
Heterogeneous bimodal attention fusion for speech emotion recognition
por: Luo, Jiachen, et al.
Publicado: (2025) -
Graph-based multi-Feature fusion method for speech emotion recognition
por: Liu, Xueyu, et al.
Publicado: (2024) -
Beyond saliency: enhancing explanation of speech emotion recognition with expert-referenced acoustic cues
por: Nasr, Seham, et al.
Publicado: (2025)