DrVoice: Parallel Speech-Text Voice Conversation Model via Dual-Resolution Speech Representations
Fuente:
arXiv
Salvato in:
| Autori principali: | Tan, Chao-Hong, Chen, Qian, Wang, Wen, Deng, Chong, Zhang, Qinglin, Cheng, Luyao, Yu, Hai, Zhang, Xin, Lv, Xiang, Zhao, Tianyu, Zhang, Chong, Ma, Yukun, Chen, Yafeng, Wang, Hui, Liu, Jiaqing, Li, Xiangang, Ye, Jieping |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
OmniFlatten: An End-to-end GPT Model for Seamless Voice Conversation
di: Zhang, Qinglin, et al.
Pubblicazione: (2024)
di: Zhang, Qinglin, et al.
Pubblicazione: (2024)
FGGM: Fisher-Guided Gradient Masking for Continual Learning
di: Tan, Chao-Hong, et al.
Pubblicazione: (2026)
di: Tan, Chao-Hong, et al.
Pubblicazione: (2026)
MAIN-VC: Lightweight Speech Representation Disentanglement for One-shot Voice Conversion
di: Li, Pengcheng, et al.
Pubblicazione: (2024)
di: Li, Pengcheng, et al.
Pubblicazione: (2024)
CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training
di: Du, Zhihao, et al.
Pubblicazione: (2025)
di: Du, Zhihao, et al.
Pubblicazione: (2025)
DreamVoice: Text-Guided Voice Conversion
di: Hai, Jiarui, et al.
Pubblicazione: (2024)
di: Hai, Jiarui, et al.
Pubblicazione: (2024)
Recording for Eyes, Not Echoing to Ears: Contextualized Spoken-to-Written Conversion of ASR Transcripts
di: Liu, Jiaqing, et al.
Pubblicazione: (2024)
di: Liu, Jiaqing, et al.
Pubblicazione: (2024)
CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models
di: Du, Zhihao, et al.
Pubblicazione: (2024)
di: Du, Zhihao, et al.
Pubblicazione: (2024)
Loss Masking Is Not Needed in Decoder-only Transformer for Discrete-token-based ASR
di: Chen, Qian, et al.
Pubblicazione: (2023)
di: Chen, Qian, et al.
Pubblicazione: (2023)
Skip-Layer Attention: Bridging Abstract and Detailed Dependencies in Transformers
di: Chen, Qian, et al.
Pubblicazione: (2024)
di: Chen, Qian, et al.
Pubblicazione: (2024)
Fun-Audio-Chat Technical Report
di: Tongyi Fun Team, et al.
Pubblicazione: (2025)
di: Tongyi Fun Team, et al.
Pubblicazione: (2025)
EmoVoice: LLM-based Emotional Text-To-Speech Model with Freestyle Text Prompting
di: Yang, Guanrou, et al.
Pubblicazione: (2025)
di: Yang, Guanrou, et al.
Pubblicazione: (2025)
Multimodal Fusion and Coherence Modeling for Video Topic Segmentation
di: Yu, Hai, et al.
Pubblicazione: (2024)
di: Yu, Hai, et al.
Pubblicazione: (2024)
VoiceX: A Text-To-Speech Framework for Custom Voices
di: Mertes, Silvan, et al.
Pubblicazione: (2024)
di: Mertes, Silvan, et al.
Pubblicazione: (2024)
MimicLM: Zero-Shot Voice Imitation through Autoregressive Modeling of Pseudo-Parallel Speech Corpora
di: Feng, Tao, et al.
Pubblicazione: (2026)
di: Feng, Tao, et al.
Pubblicazione: (2026)
EAD-VC: Enhancing Speech Auto-Disentanglement for Voice Conversion with IFUB Estimator and Joint Text-Guided Consistent Learning
di: Liang, Ziqi, et al.
Pubblicazione: (2024)
di: Liang, Ziqi, et al.
Pubblicazione: (2024)
Vevo2: A Unified and Controllable Framework for Speech and Singing Voice Generation
di: Zhang, Xueyao, et al.
Pubblicazione: (2025)
di: Zhang, Xueyao, et al.
Pubblicazione: (2025)
MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
di: Chen, Qian, et al.
Pubblicazione: (2025)
di: Chen, Qian, et al.
Pubblicazione: (2025)
Cross-lingual Text-To-Speech with Flow-based Voice Conversion for Improved Pronunciation
di: Ellinas, Nikolaos, et al.
Pubblicazione: (2022)
di: Ellinas, Nikolaos, et al.
Pubblicazione: (2022)
Self-Distillation Prototypes Network: Learning Robust Speaker Representations without Supervision
di: Chen, Yafeng, et al.
Pubblicazione: (2023)
di: Chen, Yafeng, et al.
Pubblicazione: (2023)
VoiceFlow: Efficient Text-to-Speech with Rectified Flow Matching
di: Guo, Yiwei, et al.
Pubblicazione: (2023)
di: Guo, Yiwei, et al.
Pubblicazione: (2023)
Revival with Voice: Multi-modal Controllable Text-to-Speech Synthesis
di: Kim, Minsu, et al.
Pubblicazione: (2025)
di: Kim, Minsu, et al.
Pubblicazione: (2025)
SpeakerLM: End-to-End Versatile Speaker Diarization and Recognition with Multimodal Large Language Models
di: Yin, Han, et al.
Pubblicazione: (2025)
di: Yin, Han, et al.
Pubblicazione: (2025)
Mitigating Unauthorized Speech Synthesis for Voice Protection
di: Zhang, Zhisheng, et al.
Pubblicazione: (2024)
di: Zhang, Zhisheng, et al.
Pubblicazione: (2024)
SF-Speech: Straightened Flow for Zero-Shot Voice Clone
di: Li, Xuyuan, et al.
Pubblicazione: (2024)
di: Li, Xuyuan, et al.
Pubblicazione: (2024)
Auden-Voice: General-Purpose Voice Encoder for Speech and Language Understanding
di: Huo, Mingyue, et al.
Pubblicazione: (2025)
di: Huo, Mingyue, et al.
Pubblicazione: (2025)
Voice-ENHANCE: Speech Restoration using a Diffusion-based Voice Conversion Framework
di: Byun, Kyungguen, et al.
Pubblicazione: (2025)
di: Byun, Kyungguen, et al.
Pubblicazione: (2025)
VoiceCraft: Zero-Shot Speech Editing and Text-to-Speech in the Wild
di: Peng, Puyuan, et al.
Pubblicazione: (2024)
di: Peng, Puyuan, et al.
Pubblicazione: (2024)
StarVC: A Unified Auto-Regressive Framework for Joint Text and Speech Generation in Voice Conversion
di: Li, Fengjin, et al.
Pubblicazione: (2025)
di: Li, Fengjin, et al.
Pubblicazione: (2025)
Conditional Latent Diffusion-Based Speech Enhancement Via Dual Context Learning
di: Zhao, Shengkui, et al.
Pubblicazione: (2025)
di: Zhao, Shengkui, et al.
Pubblicazione: (2025)
Speech to Speech Synthesis for Voice Impersonation
di: Johnson, Bjorn, et al.
Pubblicazione: (2026)
di: Johnson, Bjorn, et al.
Pubblicazione: (2026)
VoiceShop: A Unified Speech-to-Speech Framework for Identity-Preserving Zero-Shot Voice Editing
di: Anastassiou, Philip, et al.
Pubblicazione: (2024)
di: Anastassiou, Philip, et al.
Pubblicazione: (2024)
Connecting Voices: LoReSpeech as a Low-Resource Speech Parallel Corpus
di: Ouzerrout, Samy
Pubblicazione: (2025)
di: Ouzerrout, Samy
Pubblicazione: (2025)
Adopting Underdogs' Ideas Triggers Fairness? When and How Underachievers' Voice Endorsement Promotes Team Voice
di: Dan Ni, et al.
Pubblicazione: (2024)
di: Dan Ni, et al.
Pubblicazione: (2024)
SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant
di: Hou, Yixuan, et al.
Pubblicazione: (2025)
di: Hou, Yixuan, et al.
Pubblicazione: (2025)
EmoAttack: Utilizing Emotional Voice Conversion for Speech Backdoor Attacks on Deep Speech Classification Models
di: Yao, Wenhan, et al.
Pubblicazione: (2024)
di: Yao, Wenhan, et al.
Pubblicazione: (2024)
UniSS: Unified Expressive Speech-to-Speech Translation with Your Voice
di: Cheng, Sitong, et al.
Pubblicazione: (2025)
di: Cheng, Sitong, et al.
Pubblicazione: (2025)
Synthetic Voices, Real Threats: Evaluating Large Text-to-Speech Models in Generating Harmful Audio
di: Chen, Guangke, et al.
Pubblicazione: (2025)
di: Chen, Guangke, et al.
Pubblicazione: (2025)
REWIND: Speech Time Reversal for Enhancing Speaker Representations in Diffusion-based Voice Conversion
di: Biyani, Ishan D., et al.
Pubblicazione: (2025)
di: Biyani, Ishan D., et al.
Pubblicazione: (2025)
SpikeVoice: High-Quality Text-to-Speech Via Efficient Spiking Neural Network
di: Wang, Kexin, et al.
Pubblicazione: (2024)
di: Wang, Kexin, et al.
Pubblicazione: (2024)
MambaVoiceCloning: Efficient and Expressive Text-to-Speech via State-Space Modeling and Diffusion Control
di: Kumar, Sahil, et al.
Pubblicazione: (2026)
di: Kumar, Sahil, et al.
Pubblicazione: (2026)
Documenti analoghi
-
OmniFlatten: An End-to-end GPT Model for Seamless Voice Conversation
di: Zhang, Qinglin, et al.
Pubblicazione: (2024) -
FGGM: Fisher-Guided Gradient Masking for Continual Learning
di: Tan, Chao-Hong, et al.
Pubblicazione: (2026) -
MAIN-VC: Lightweight Speech Representation Disentanglement for One-shot Voice Conversion
di: Li, Pengcheng, et al.
Pubblicazione: (2024) -
CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training
di: Du, Zhihao, et al.
Pubblicazione: (2025) -
DreamVoice: Text-Guided Voice Conversion
di: Hai, Jiarui, et al.
Pubblicazione: (2024)