Serial-Parallel Dual-Path Architecture for Speaking Style Recognition
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Guojian, Shao, Qijie, Zhao, Zhixian, Wang, Shuiyuan, Fu, Zhonghua, Xie, Lei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
FaceSpeak: Expressive and High-Quality Speech Synthesis from Human Portraits of Different Styles
von: Zhang, Tian-Hao, et al.
Veröffentlicht: (2025)
von: Zhang, Tian-Hao, et al.
Veröffentlicht: (2025)
Steering Language Model to Stable Speech Emotion Recognition via Contextual Perception and Chain of Thought
von: Zhao, Zhixian, et al.
Veröffentlicht: (2025)
von: Zhao, Zhixian, et al.
Veröffentlicht: (2025)
ReStyle-TTS: Relative and Continuous Style Control for Zero-Shot Speech Synthesis
von: Li, Haitao, et al.
Veröffentlicht: (2026)
von: Li, Haitao, et al.
Veröffentlicht: (2026)
HumDial-EIBench: A Human-Recorded Multi-Turn Emotional Intelligence Benchmark for Audio Language Models
von: Wang, Shuiyuan, et al.
Veröffentlicht: (2026)
von: Wang, Shuiyuan, et al.
Veröffentlicht: (2026)
Serialized Speech Information Guidance with Overlapped Encoding Separation for Multi-Speaker Automatic Speech Recognition
von: Shi, Hao, et al.
Veröffentlicht: (2024)
von: Shi, Hao, et al.
Veröffentlicht: (2024)
Serialized Output Training by Learned Dominance
von: Shi, Ying, et al.
Veröffentlicht: (2024)
von: Shi, Ying, et al.
Veröffentlicht: (2024)
DualSpec: Text-to-spatial-audio Generation via Dual-Spectrogram Guided Diffusion Model
von: Zhao, Lei, et al.
Veröffentlicht: (2025)
von: Zhao, Lei, et al.
Veröffentlicht: (2025)
Text Prompt is Not Enough: Sound Event Enhanced Prompt Adapter for Target Style Audio Generation
von: Xiong, Chenxu, et al.
Veröffentlicht: (2024)
von: Xiong, Chenxu, et al.
Veröffentlicht: (2024)
X-Voice: Enabling Everyone to Speak 30 Languages via Zero-Shot Cross-Lingual Voice Cloning
von: Xu, Rixi, et al.
Veröffentlicht: (2026)
von: Xu, Rixi, et al.
Veröffentlicht: (2026)
An Effective Automated Speaking Assessment Approach to Mitigating Data Scarcity and Imbalanced Distribution
von: Lo, Tien-Hong, et al.
Veröffentlicht: (2024)
von: Lo, Tien-Hong, et al.
Veröffentlicht: (2024)
Disentangled Dual-Branch Graph Learning for Conversational Emotion Recognition
von: Guo, Chengling, et al.
Veröffentlicht: (2026)
von: Guo, Chengling, et al.
Veröffentlicht: (2026)
MLCA-AVSR: Multi-Layer Cross Attention Fusion based Audio-Visual Speech Recognition
von: Wang, He, et al.
Veröffentlicht: (2024)
von: Wang, He, et al.
Veröffentlicht: (2024)
DPI-TTS: Directional Patch Interaction for Fast-Converging and Style Temporal Modeling in Text-to-Speech
von: Qi, Xin, et al.
Veröffentlicht: (2024)
von: Qi, Xin, et al.
Veröffentlicht: (2024)
Automatic Speech Recognition in the Modern Era: Architectures, Training, and Evaluation
von: Nayeem, Md., et al.
Veröffentlicht: (2025)
von: Nayeem, Md., et al.
Veröffentlicht: (2025)
Meta-Learning Empowered Meta-Face: Personalized Speaking Style Adaptation for Audio-Driven 3D Talking Face Animation
von: Zhou, Xukun, et al.
Veröffentlicht: (2024)
von: Zhou, Xukun, et al.
Veröffentlicht: (2024)
Speech-to-Speech Translation with Discrete-Unit-Based Style Transfer
von: Wang, Yongqi, et al.
Veröffentlicht: (2023)
von: Wang, Yongqi, et al.
Veröffentlicht: (2023)
The NPU-ASLP-LiAuto System Description for Visual Speech Recognition in CNVSRC 2023
von: Wang, He, et al.
Veröffentlicht: (2024)
von: Wang, He, et al.
Veröffentlicht: (2024)
Leveraging Label Potential for Enhanced Multimodal Emotion Recognition
von: Shao, Xuechun, et al.
Veröffentlicht: (2025)
von: Shao, Xuechun, et al.
Veröffentlicht: (2025)
dLLM-ASR: A Faster Diffusion LLM-based Framework for Speech Recognition
von: Tian, Wenjie, et al.
Veröffentlicht: (2026)
von: Tian, Wenjie, et al.
Veröffentlicht: (2026)
Spotlight-TTS: Spotlighting the Style via Voiced-Aware Style Extraction and Style Direction Adjustment for Expressive Text-to-Speech
von: Kim, Nam-Gyu, et al.
Veröffentlicht: (2025)
von: Kim, Nam-Gyu, et al.
Veröffentlicht: (2025)
MMGER: Multi-modal and Multi-granularity Generative Error Correction with LLM for Joint Accent and Speech Recognition
von: Mu, Bingshen, et al.
Veröffentlicht: (2024)
von: Mu, Bingshen, et al.
Veröffentlicht: (2024)
Convert and Speak: Zero-shot Accent Conversion with Minimum Supervision
von: Jia, Zhijun, et al.
Veröffentlicht: (2024)
von: Jia, Zhijun, et al.
Veröffentlicht: (2024)
GMP-TL: Gender-augmented Multi-scale Pseudo-label Enhanced Transfer Learning for Speech Emotion Recognition
von: Pan, Yu, et al.
Veröffentlicht: (2024)
von: Pan, Yu, et al.
Veröffentlicht: (2024)
Improvement and Implementation of a Speech Emotion Recognition Model Based on Dual-Layer LSTM
von: Yang, Xiaoran, et al.
Veröffentlicht: (2024)
von: Yang, Xiaoran, et al.
Veröffentlicht: (2024)
HDMoLE: Mixture of LoRA Experts with Hierarchical Routing and Dynamic Thresholds for Fine-Tuning LLM-based ASR Models
von: Mu, Bingshen, et al.
Veröffentlicht: (2024)
von: Mu, Bingshen, et al.
Veröffentlicht: (2024)
Searching for Effective Preprocessing Method and CNN-based Architecture with Efficient Channel Attention on Speech Emotion Recognition
von: Kim, Byunggun, et al.
Veröffentlicht: (2024)
von: Kim, Byunggun, et al.
Veröffentlicht: (2024)
DSFlow: Dual Supervision and Step-Aware Architecture for One-Step Flow Matching Speech Synthesis
von: Lin, Bin, et al.
Veröffentlicht: (2026)
von: Lin, Bin, et al.
Veröffentlicht: (2026)
Unifying Speech Recognition, Synthesis and Conversion with Autoregressive Transformers
von: Cai, Runyuan, et al.
Veröffentlicht: (2026)
von: Cai, Runyuan, et al.
Veröffentlicht: (2026)
Ideal-LLM: Integrating Dual Encoders and Language-Adapted LLM for Multilingual Speech-to-Text
von: Xue, Hongfei, et al.
Veröffentlicht: (2024)
von: Xue, Hongfei, et al.
Veröffentlicht: (2024)
Magnitude-Phase Dual-Path Speech Enhancement Network based on Self-Supervised Embedding and Perceptual Contrast Stretch Boosting
von: Mattursun, Alimjan, et al.
Veröffentlicht: (2025)
von: Mattursun, Alimjan, et al.
Veröffentlicht: (2025)
DQ-Data2vec: Decoupling Quantization for Multilingual Speech Recognition
von: Shao, Qijie, et al.
Veröffentlicht: (2025)
von: Shao, Qijie, et al.
Veröffentlicht: (2025)
ICMC-ASR: The ICASSP 2024 In-Car Multi-Channel Automatic Speech Recognition Challenge
von: Wang, He, et al.
Veröffentlicht: (2024)
von: Wang, He, et al.
Veröffentlicht: (2024)
CAMEL: Cross-Attention Enhanced Mixture-of-Experts and Language Bias for Code-Switching Speech Recognition
von: Wang, He, et al.
Veröffentlicht: (2024)
von: Wang, He, et al.
Veröffentlicht: (2024)
OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia
von: Geng, Xuelong, et al.
Veröffentlicht: (2025)
von: Geng, Xuelong, et al.
Veröffentlicht: (2025)
Self-supervised ASR Models and Features For Dysarthric and Elderly Speech Recognition
von: Hu, Shujie, et al.
Veröffentlicht: (2024)
von: Hu, Shujie, et al.
Veröffentlicht: (2024)
Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
von: Wang, Xinsheng, et al.
Veröffentlicht: (2025)
von: Wang, Xinsheng, et al.
Veröffentlicht: (2025)
ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech
von: Pan, Yu, et al.
Veröffentlicht: (2025)
von: Pan, Yu, et al.
Veröffentlicht: (2025)
Breaking Resource Barriers in Speech Emotion Recognition via Data Distillation
von: Chang, Yi, et al.
Veröffentlicht: (2024)
von: Chang, Yi, et al.
Veröffentlicht: (2024)
Seeing the Context: Rich Visual Context-Aware Speech Recognition via Multimodal Reasoning
von: Tian, Wenjie, et al.
Veröffentlicht: (2026)
von: Tian, Wenjie, et al.
Veröffentlicht: (2026)
Serialized Output Prompting for Large Language Model-based Multi-Talker Speech Recognition
von: Shi, Hao, et al.
Veröffentlicht: (2025)
von: Shi, Hao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
FaceSpeak: Expressive and High-Quality Speech Synthesis from Human Portraits of Different Styles
von: Zhang, Tian-Hao, et al.
Veröffentlicht: (2025) -
Steering Language Model to Stable Speech Emotion Recognition via Contextual Perception and Chain of Thought
von: Zhao, Zhixian, et al.
Veröffentlicht: (2025) -
ReStyle-TTS: Relative and Continuous Style Control for Zero-Shot Speech Synthesis
von: Li, Haitao, et al.
Veröffentlicht: (2026) -
HumDial-EIBench: A Human-Recorded Multi-Turn Emotional Intelligence Benchmark for Audio Language Models
von: Wang, Shuiyuan, et al.
Veröffentlicht: (2026) -
Serialized Speech Information Guidance with Overlapped Encoding Separation for Multi-Speaker Automatic Speech Recognition
von: Shi, Hao, et al.
Veröffentlicht: (2024)