SemaVoice: Semantic-Aware Continuous Autoregressive Speech Synthesis
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Huimeng, Lu, Hui, Deng, Jiajun, Xu, Haoning, Chen, Youjun, Chen, Xueyuan, Li, Zhaoqing, Peng, Shuhai, Kang, Shiyin, Liu, Xunying |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Phone-purity Guided Discrete Tokens for Dysarthric Speech Recognition
von: Wang, Huimeng, et al.
Veröffentlicht: (2025)
von: Wang, Huimeng, et al.
Veröffentlicht: (2025)
Effective and Efficient One-pass Compression of Speech Foundation Models Using Sparsity-aware Self-pinching Gates
von: Xu, Haoning, et al.
Veröffentlicht: (2025)
von: Xu, Haoning, et al.
Veröffentlicht: (2025)
Effective and Efficient Mixed Precision Quantization of Speech Foundation Models
von: Xu, Haoning, et al.
Veröffentlicht: (2025)
von: Xu, Haoning, et al.
Veröffentlicht: (2025)
How Should LLMs Listen While Speaking? A Study of User-Stream Routing in Full-Duplex Spoken Dialogue
von: Lu, Hui, et al.
Veröffentlicht: (2026)
von: Lu, Hui, et al.
Veröffentlicht: (2026)
MOPSA: Mixture of Prompt-Experts Based Speaker Adaptation for Elderly Speech Recognition
von: Deng, Chengxi, et al.
Veröffentlicht: (2025)
von: Deng, Chengxi, et al.
Veröffentlicht: (2025)
Towards One-bit ASR: Extremely Low-bit Conformer Quantization Using Co-training and Stochastic Precision
von: Li, Zhaoqing, et al.
Veröffentlicht: (2025)
von: Li, Zhaoqing, et al.
Veröffentlicht: (2025)
Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition
von: Chen, Youjun, et al.
Veröffentlicht: (2025)
von: Chen, Youjun, et al.
Veröffentlicht: (2025)
Unfolding A Few Structures for The Many: Memory-Efficient Compression of Conformer and Speech Foundation Models
von: Li, Zhaoqing, et al.
Veröffentlicht: (2025)
von: Li, Zhaoqing, et al.
Veröffentlicht: (2025)
Exploring Cross-Utterance Speech Contexts for Conformer-Transducer Speech Recognition Systems
von: Cui, Mingyu, et al.
Veröffentlicht: (2025)
von: Cui, Mingyu, et al.
Veröffentlicht: (2025)
UNISON: A Unified Sound Generation and Editing Framework via Deep LLM Fusion
von: Li, Zhaoqing, et al.
Veröffentlicht: (2026)
von: Li, Zhaoqing, et al.
Veröffentlicht: (2026)
Enhancing Pre-trained ASR System Fine-tuning for Dysarthric Speech Recognition using Adversarial Data Augmentation
von: Wang, Huimeng, et al.
Veröffentlicht: (2024)
von: Wang, Huimeng, et al.
Veröffentlicht: (2024)
CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis
von: Wu, Chun Yat, et al.
Veröffentlicht: (2025)
von: Wu, Chun Yat, et al.
Veröffentlicht: (2025)
On-the-fly Routing for Zero-shot MoE Speaker Adaptation of Speech Foundation Models for Dysarthric Speech Recognition
von: HU, Shujie, et al.
Veröffentlicht: (2025)
von: HU, Shujie, et al.
Veröffentlicht: (2025)
One-pass Multiple Conformer and Foundation Speech Systems Compression and Quantization Using An All-in-one Neural Model
von: Li, Zhaoqing, et al.
Veröffentlicht: (2024)
von: Li, Zhaoqing, et al.
Veröffentlicht: (2024)
Exploring SSL Discrete Speech Features for Zipformer-based Contextual ASR
von: Cui, Mingyu, et al.
Veröffentlicht: (2024)
von: Cui, Mingyu, et al.
Veröffentlicht: (2024)
Joint Speaker Features Learning for Audio-visual Multichannel Speech Separation and Recognition
von: Li, Guinan, et al.
Veröffentlicht: (2024)
von: Li, Guinan, et al.
Veröffentlicht: (2024)
Efficient Adapter Tuning for Joint Singing Voice Beat and Downbeat Tracking with Self-supervised Learning Features
von: Deng, Jiajun, et al.
Veröffentlicht: (2025)
von: Deng, Jiajun, et al.
Veröffentlicht: (2025)
Structured Speaker-Deficiency Adaptation of Foundation Models for Dysarthric and Elderly Speech Recognition
von: Hu, Shujie, et al.
Veröffentlicht: (2024)
von: Hu, Shujie, et al.
Veröffentlicht: (2024)
Exploiting Audio-Visual Features with Pretrained AV-HuBERT for Multi-Modal Dysarthric Speech Reconstruction
von: Chen, Xueyuan, et al.
Veröffentlicht: (2024)
von: Chen, Xueyuan, et al.
Veröffentlicht: (2024)
Personalized Adversarial Data Augmentation for Dysarthric and Elderly Speech Recognition
von: Jin, Zengrui, et al.
Veröffentlicht: (2022)
von: Jin, Zengrui, et al.
Veröffentlicht: (2022)
ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching
von: Zhu, Han, et al.
Veröffentlicht: (2025)
von: Zhu, Han, et al.
Veröffentlicht: (2025)
Semantic-VAE: Semantic-Alignment Latent Representation for Better Speech Synthesis
von: Niu, Zhikang, et al.
Veröffentlicht: (2025)
von: Niu, Zhikang, et al.
Veröffentlicht: (2025)
Towards Effective and Efficient Non-autoregressive decoders for Conformer and LLM-based ASR using Block-based Attention Mask
von: Wang, Tianzi, et al.
Veröffentlicht: (2025)
von: Wang, Tianzi, et al.
Veröffentlicht: (2025)
Parallel Synthesis for Autoregressive Speech Generation
von: Hsu, Po-chun, et al.
Veröffentlicht: (2022)
von: Hsu, Po-chun, et al.
Veröffentlicht: (2022)
Incremental Disentanglement for Environment-Aware Zero-Shot Text-to-Speech Synthesis
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2024)
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2024)
Exploring SSL Discrete Tokens for Multilingual ASR
von: Cui, Mingyu, et al.
Veröffentlicht: (2024)
von: Cui, Mingyu, et al.
Veröffentlicht: (2024)
ZipVoice-Dialog: Non-Autoregressive Spoken Dialogue Generation with Flow Matching
von: Zhu, Han, et al.
Veröffentlicht: (2025)
von: Zhu, Han, et al.
Veröffentlicht: (2025)
VoiceDiT: Dual-Condition Diffusion Transformer for Environment-Aware Speech Synthesis
von: Jung, Jaemin, et al.
Veröffentlicht: (2024)
von: Jung, Jaemin, et al.
Veröffentlicht: (2024)
Disentangling Speakers in Multi-Talker Speech Recognition with Speaker-Aware CTC
von: Kang, Jiawen, et al.
Veröffentlicht: (2024)
von: Kang, Jiawen, et al.
Veröffentlicht: (2024)
Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis
von: Lin, Weiwei, et al.
Veröffentlicht: (2025)
von: Lin, Weiwei, et al.
Veröffentlicht: (2025)
Variational Auto-Encoder Based Variability Encoding for Dysarthric Speech Recognition
von: Xie, Xurong, et al.
Veröffentlicht: (2022)
von: Xie, Xurong, et al.
Veröffentlicht: (2022)
VoiceCraft-X: Unifying Multilingual, Voice-Cloning Speech Synthesis and Speech Editing
von: Zheng, Zhisheng, et al.
Veröffentlicht: (2025)
von: Zheng, Zhisheng, et al.
Veröffentlicht: (2025)
Direct Preference Optimization for Speech Autoregressive Diffusion Models
von: Liu, Zhijun, et al.
Veröffentlicht: (2025)
von: Liu, Zhijun, et al.
Veröffentlicht: (2025)
Revival with Voice: Multi-modal Controllable Text-to-Speech Synthesis
von: Kim, Minsu, et al.
Veröffentlicht: (2025)
von: Kim, Minsu, et al.
Veröffentlicht: (2025)
Investigation of Deep Neural Network Acoustic Modelling Approaches for Low Resource Accented Mandarin Speech Recognition
von: Xie, Xurong, et al.
Veröffentlicht: (2022)
von: Xie, Xurong, et al.
Veröffentlicht: (2022)
Improving Language Model-Based Zero-Shot Text-to-Speech Synthesis with Multi-Scale Acoustic Prompts
von: Lei, Shun, et al.
Veröffentlicht: (2023)
von: Lei, Shun, et al.
Veröffentlicht: (2023)
Speech to Speech Synthesis for Voice Impersonation
von: Johnson, Bjorn, et al.
Veröffentlicht: (2026)
von: Johnson, Bjorn, et al.
Veröffentlicht: (2026)
A Comparative Study of Discrete Speech Tokens for Semantic-Related Tasks with Large Language Models
von: Wang, Dingdong, et al.
Veröffentlicht: (2024)
von: Wang, Dingdong, et al.
Veröffentlicht: (2024)
Speech Synthesis along Perceptual Voice Quality Dimensions
von: Rautenberg, Frederik, et al.
Veröffentlicht: (2025)
von: Rautenberg, Frederik, et al.
Veröffentlicht: (2025)
Autoregressive Speech Synthesis without Vector Quantization
von: Meng, Lingwei, et al.
Veröffentlicht: (2024)
von: Meng, Lingwei, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Phone-purity Guided Discrete Tokens for Dysarthric Speech Recognition
von: Wang, Huimeng, et al.
Veröffentlicht: (2025) -
Effective and Efficient One-pass Compression of Speech Foundation Models Using Sparsity-aware Self-pinching Gates
von: Xu, Haoning, et al.
Veröffentlicht: (2025) -
Effective and Efficient Mixed Precision Quantization of Speech Foundation Models
von: Xu, Haoning, et al.
Veröffentlicht: (2025) -
How Should LLMs Listen While Speaking? A Study of User-Stream Routing in Full-Duplex Spoken Dialogue
von: Lu, Hui, et al.
Veröffentlicht: (2026) -
MOPSA: Mixture of Prompt-Experts Based Speaker Adaptation for Elderly Speech Recognition
von: Deng, Chengxi, et al.
Veröffentlicht: (2025)