VoiceBridge: General Speech Restoration with One-step Latent Bridge Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhang, Chi, Zheng, Kaiwen, Chen, Zehua, Zhu, Jun |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Schrodinger Bridges Beat Diffusion Models on Text-to-Speech Synthesis
di: Chen, Zehua, et al.
Pubblicazione: (2023)
di: Chen, Zehua, et al.
Pubblicazione: (2023)
Bridge-SR: Schrödinger Bridge for Efficient SR
di: Li, Chang, et al.
Pubblicazione: (2025)
di: Li, Chang, et al.
Pubblicazione: (2025)
WaveLLDM: Design and Development of a Lightweight Latent Diffusion Model for Speech Enhancement and Restoration
di: Santoso, Kevin Putra, et al.
Pubblicazione: (2025)
di: Santoso, Kevin Putra, et al.
Pubblicazione: (2025)
Tiny-Align: Bridging Automatic Speech Recognition and Large Language Model on the Edge
di: Qin, Ruiyang, et al.
Pubblicazione: (2024)
di: Qin, Ruiyang, et al.
Pubblicazione: (2024)
Bridging ASR and LLMs for Dysarthric Speech Recognition: Benchmarking Self-Supervised and Generative Approaches
di: Aboeitta, Ahmed, et al.
Pubblicazione: (2025)
di: Aboeitta, Ahmed, et al.
Pubblicazione: (2025)
Schrödinger Bridge Mamba for One-Step Speech Enhancement
di: Yang, Jing, et al.
Pubblicazione: (2025)
di: Yang, Jing, et al.
Pubblicazione: (2025)
Few-step Adversarial Schrödinger Bridge for Generative Speech Enhancement
di: Han, Seungu, et al.
Pubblicazione: (2025)
di: Han, Seungu, et al.
Pubblicazione: (2025)
Interpreting Pretrained Speech Models for Automatic Speech Assessment of Voice Disorders
di: Lau, Hok-Shing, et al.
Pubblicazione: (2024)
di: Lau, Hok-Shing, et al.
Pubblicazione: (2024)
OZSpeech: One-step Zero-shot Speech Synthesis with Learned-Prior-Conditioned Flow Matching
di: Huynh-Nguyen, Hieu-Nghia, et al.
Pubblicazione: (2025)
di: Huynh-Nguyen, Hieu-Nghia, et al.
Pubblicazione: (2025)
High-Resolution Speech Restoration with Latent Diffusion Model
di: Dhyani, Tushar, et al.
Pubblicazione: (2024)
di: Dhyani, Tushar, et al.
Pubblicazione: (2024)
EmoAttack: Utilizing Emotional Voice Conversion for Speech Backdoor Attacks on Deep Speech Classification Models
di: Yao, Wenhan, et al.
Pubblicazione: (2024)
di: Yao, Wenhan, et al.
Pubblicazione: (2024)
CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training
di: Du, Zhihao, et al.
Pubblicazione: (2025)
di: Du, Zhihao, et al.
Pubblicazione: (2025)
MeanVoiceFlow: One-step Nonparallel Voice Conversion with Mean Flows
di: Kaneko, Takuhiro, et al.
Pubblicazione: (2026)
di: Kaneko, Takuhiro, et al.
Pubblicazione: (2026)
VoiceShop: A Unified Speech-to-Speech Framework for Identity-Preserving Zero-Shot Voice Editing
di: Anastassiou, Philip, et al.
Pubblicazione: (2024)
di: Anastassiou, Philip, et al.
Pubblicazione: (2024)
DeepASMR: LLM-Based Zero-Shot ASMR Speech Generation for Anyone of Any Voice
di: Zhang, Leying, et al.
Pubblicazione: (2026)
di: Zhang, Leying, et al.
Pubblicazione: (2026)
Emotion-Aware Speech Generation with Character-Specific Voices for Comics
di: Qian, Zhiwen, et al.
Pubblicazione: (2025)
di: Qian, Zhiwen, et al.
Pubblicazione: (2025)
Voice Cloning for Dysarthric Speech Synthesis: Addressing Data Scarcity in Speech-Language Pathology
di: Moell, Birger, et al.
Pubblicazione: (2025)
di: Moell, Birger, et al.
Pubblicazione: (2025)
FastVoiceGrad: One-step Diffusion-Based Voice Conversion with Adversarial Conditional Diffusion Distillation
di: Kaneko, Takuhiro, et al.
Pubblicazione: (2024)
di: Kaneko, Takuhiro, et al.
Pubblicazione: (2024)
Generating Moving 3D Soundscapes with Latent Diffusion Models
di: Templin, Christian, et al.
Pubblicazione: (2025)
di: Templin, Christian, et al.
Pubblicazione: (2025)
CORD: Bridging the Audio-Text Reasoning Gap via Weighted On-policy Cross-modal Distillation
di: Hu, Jing, et al.
Pubblicazione: (2026)
di: Hu, Jing, et al.
Pubblicazione: (2026)
FasterVoiceGrad: Faster One-step Diffusion-Based Voice Conversion with Adversarial Diffusion Conversion Distillation
di: Kaneko, Takuhiro, et al.
Pubblicazione: (2025)
di: Kaneko, Takuhiro, et al.
Pubblicazione: (2025)
Improving Voice Quality in Speech Anonymization With Just Perception-Informed Losses
di: Ghosh, Suhita, et al.
Pubblicazione: (2024)
di: Ghosh, Suhita, et al.
Pubblicazione: (2024)
Speech Foundation Model Ensembles for the Controlled Singing Voice Deepfake Detection (CtrSVDD) Challenge 2024
di: Guragain, Anmol, et al.
Pubblicazione: (2024)
di: Guragain, Anmol, et al.
Pubblicazione: (2024)
FreeAudio: Training-Free Timing Planning for Controllable Long-Form Text-to-Audio Generation
di: Jiang, Yuxuan, et al.
Pubblicazione: (2025)
di: Jiang, Yuxuan, et al.
Pubblicazione: (2025)
Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
di: Wang, Xinsheng, et al.
Pubblicazione: (2025)
di: Wang, Xinsheng, et al.
Pubblicazione: (2025)
AnyEnhance: A Unified Generative Model with Prompt-Guidance and Self-Critic for Voice Enhancement
di: Zhang, Junan, et al.
Pubblicazione: (2025)
di: Zhang, Junan, et al.
Pubblicazione: (2025)
FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs
di: An, Keyu, et al.
Pubblicazione: (2024)
di: An, Keyu, et al.
Pubblicazione: (2024)
ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment
di: Zhao, Shengkui, et al.
Pubblicazione: (2025)
di: Zhao, Shengkui, et al.
Pubblicazione: (2025)
Multi-Metric Preference Alignment for Generative Speech Restoration
di: Zhang, Junan, et al.
Pubblicazione: (2025)
di: Zhang, Junan, et al.
Pubblicazione: (2025)
Neural Speech Embeddings for Speech Synthesis Based on Deep Generative Networks
di: Lee, Seo-Hyun, et al.
Pubblicazione: (2023)
di: Lee, Seo-Hyun, et al.
Pubblicazione: (2023)
DiffGAP: A Lightweight Diffusion Module in Contrastive Space for Bridging Cross-Model Gap
di: Mo, Shentong, et al.
Pubblicazione: (2025)
di: Mo, Shentong, et al.
Pubblicazione: (2025)
Do Not Mimic My Voice: Speaker Identity Unlearning for Zero-Shot Text-to-Speech
di: Kim, Taesoo, et al.
Pubblicazione: (2025)
di: Kim, Taesoo, et al.
Pubblicazione: (2025)
Defense Against Synthetic Speech: Real-Time Detection of RVC Voice Conversion Attacks
di: Chinchmalatpure, Prajwal, et al.
Pubblicazione: (2025)
di: Chinchmalatpure, Prajwal, et al.
Pubblicazione: (2025)
Latent Diffusion Bridges for Unsupervised Musical Audio Timbre Transfer
di: Mancusi, Michele, et al.
Pubblicazione: (2024)
di: Mancusi, Michele, et al.
Pubblicazione: (2024)
Mitigating Latent Mismatch in cVAE-Based Singing Voice Synthesis via Flow Matching
di: Yun, Minhyeok, et al.
Pubblicazione: (2026)
di: Yun, Minhyeok, et al.
Pubblicazione: (2026)
LHQ-SVC: Lightweight and High Quality Singing Voice Conversion Modeling
di: Huang, Yubo, et al.
Pubblicazione: (2024)
di: Huang, Yubo, et al.
Pubblicazione: (2024)
Pureformer-VC: Non-parallel One-Shot Voice Conversion with Pure Transformer Blocks and Triplet Discriminative Training
di: Yao, Wenhan, et al.
Pubblicazione: (2024)
di: Yao, Wenhan, et al.
Pubblicazione: (2024)
Effective and Efficient One-pass Compression of Speech Foundation Models Using Sparsity-aware Self-pinching Gates
di: Xu, Haoning, et al.
Pubblicazione: (2025)
di: Xu, Haoning, et al.
Pubblicazione: (2025)
DRCap: Decoding CLAP Latents with Retrieval-Augmented Generation for Zero-shot Audio Captioning
di: Li, Xiquan, et al.
Pubblicazione: (2024)
di: Li, Xiquan, et al.
Pubblicazione: (2024)
NarraScore: Bridging Visual Narrative and Musical Dynamics via Hierarchical Affective Control
di: Wen, Yufan, et al.
Pubblicazione: (2026)
di: Wen, Yufan, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Schrodinger Bridges Beat Diffusion Models on Text-to-Speech Synthesis
di: Chen, Zehua, et al.
Pubblicazione: (2023) -
Bridge-SR: Schrödinger Bridge for Efficient SR
di: Li, Chang, et al.
Pubblicazione: (2025) -
WaveLLDM: Design and Development of a Lightweight Latent Diffusion Model for Speech Enhancement and Restoration
di: Santoso, Kevin Putra, et al.
Pubblicazione: (2025) -
Tiny-Align: Bridging Automatic Speech Recognition and Large Language Model on the Edge
di: Qin, Ruiyang, et al.
Pubblicazione: (2024) -
Bridging ASR and LLMs for Dysarthric Speech Recognition: Benchmarking Self-Supervised and Generative Approaches
di: Aboeitta, Ahmed, et al.
Pubblicazione: (2025)