A Lightweight Pipeline for Noisy Speech Voice Cloning and Accurate Lip Sync Synthesis
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Amir, Javeria, Attaria, Farwa, Jabeen, Mah, Noor, Umara, Rashid, Zahid |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Voice Cloning for Dysarthric Speech Synthesis: Addressing Data Scarcity in Speech-Language Pathology
von: Moell, Birger, et al.
Veröffentlicht: (2025)
von: Moell, Birger, et al.
Veröffentlicht: (2025)
Pronunciation Deviation Analysis Through Voice Cloning and Acoustic Comparison
von: Valdivia, Andrew, et al.
Veröffentlicht: (2025)
von: Valdivia, Andrew, et al.
Veröffentlicht: (2025)
SyncSpeech: Efficient and Low-Latency Text-to-Speech based on Temporal Masked Transformer
von: Sheng, Zhengyan, et al.
Veröffentlicht: (2025)
von: Sheng, Zhengyan, et al.
Veröffentlicht: (2025)
Voice Cloning: Comprehensive Survey
von: Azzuni, Hussam, et al.
Veröffentlicht: (2025)
von: Azzuni, Hussam, et al.
Veröffentlicht: (2025)
Neural Multi-Speaker Voice Cloning for Nepali in Low-Resource Settings
von: Shrestha, Aayush M., et al.
Veröffentlicht: (2026)
von: Shrestha, Aayush M., et al.
Veröffentlicht: (2026)
Proactive Detection of Voice Cloning with Localized Watermarking
von: Roman, Robin San, et al.
Veröffentlicht: (2024)
von: Roman, Robin San, et al.
Veröffentlicht: (2024)
CloneShield: A Framework for Universal Perturbation Against Zero-Shot Voice Cloning
von: Li, Renyuan, et al.
Veröffentlicht: (2025)
von: Li, Renyuan, et al.
Veröffentlicht: (2025)
VividVoice: A Unified Framework for Scene-Aware Visually-Driven Speech Synthesis
von: Ma, Chengyuan, et al.
Veröffentlicht: (2026)
von: Ma, Chengyuan, et al.
Veröffentlicht: (2026)
The ISCSLP 2024 Conversational Voice Clone (CoVoC) Challenge: Tasks, Results and Findings
von: Xia, Kangxiang, et al.
Veröffentlicht: (2024)
von: Xia, Kangxiang, et al.
Veröffentlicht: (2024)
Voice "Cloning" is Style Transfer
von: Zhou, Kaitlyn, et al.
Veröffentlicht: (2026)
von: Zhou, Kaitlyn, et al.
Veröffentlicht: (2026)
Empowering Global Voices: A Data-Efficient, Phoneme-Tone Adaptive Approach to High-Fidelity Speech Synthesis
von: Geng, Yizhong, et al.
Veröffentlicht: (2025)
von: Geng, Yizhong, et al.
Veröffentlicht: (2025)
UniSS: Unified Expressive Speech-to-Speech Translation with Your Voice
von: Cheng, Sitong, et al.
Veröffentlicht: (2025)
von: Cheng, Sitong, et al.
Veröffentlicht: (2025)
SlimSpeech: Lightweight and Efficient Text-to-Speech with Slim Rectified Flow
von: Wang, Kaidi, et al.
Veröffentlicht: (2025)
von: Wang, Kaidi, et al.
Veröffentlicht: (2025)
X-Voice: Enabling Everyone to Speak 30 Languages via Zero-Shot Cross-Lingual Voice Cloning
von: Xu, Rixi, et al.
Veröffentlicht: (2026)
von: Xu, Rixi, et al.
Veröffentlicht: (2026)
Fed-PISA: Federated Voice Cloning via Personalized Identity-Style Adaptation
von: Wang, Qi, et al.
Veröffentlicht: (2025)
von: Wang, Qi, et al.
Veröffentlicht: (2025)
SafeSpeech: Robust and Universal Voice Protection Against Malicious Speech Synthesis
von: Zhang, Zhisheng, et al.
Veröffentlicht: (2025)
von: Zhang, Zhisheng, et al.
Veröffentlicht: (2025)
Mitigating Unauthorized Speech Synthesis for Voice Protection
von: Zhang, Zhisheng, et al.
Veröffentlicht: (2024)
von: Zhang, Zhisheng, et al.
Veröffentlicht: (2024)
SyncTrack: Rhythmic Stability and Synchronization in Multi-Track Music Generation
von: Wang, Hongrui, et al.
Veröffentlicht: (2026)
von: Wang, Hongrui, et al.
Veröffentlicht: (2026)
Voiced-Aware Style Extraction and Style Direction Adjustment for Expressive Text-to-Speech
von: Kim, Nam-Gyu
Veröffentlicht: (2025)
von: Kim, Nam-Gyu
Veröffentlicht: (2025)
Exploring Phonetic Context-Aware Lip-Sync For Talking Face Generation
von: Park, Se Jin, et al.
Veröffentlicht: (2023)
von: Park, Se Jin, et al.
Veröffentlicht: (2023)
Emotion Detection in Speech Using Lightweight and Transformer-Based Models: A Comparative and Ablation Study
von: Onyekwelu-Udoka, Lucky, et al.
Veröffentlicht: (2025)
von: Onyekwelu-Udoka, Lucky, et al.
Veröffentlicht: (2025)
VoiceShop: A Unified Speech-to-Speech Framework for Identity-Preserving Zero-Shot Voice Editing
von: Anastassiou, Philip, et al.
Veröffentlicht: (2024)
von: Anastassiou, Philip, et al.
Veröffentlicht: (2024)
MindVoice: Reconstructing Intelligible Speech from Non-invasive Neural Signals with Pretrained Priors
von: Bao, Guangyin, et al.
Veröffentlicht: (2026)
von: Bao, Guangyin, et al.
Veröffentlicht: (2026)
LSZone: A Lightweight Spatial Information Modeling Architecture for Real-time In-car Multi-zone Speech Separation
von: Chen, Jun, et al.
Veröffentlicht: (2025)
von: Chen, Jun, et al.
Veröffentlicht: (2025)
VoiceMark: Zero-Shot Voice Cloning-Resistant Watermarking Approach Leveraging Speaker-Specific Latents
von: Li, Haiyun, et al.
Veröffentlicht: (2025)
von: Li, Haiyun, et al.
Veröffentlicht: (2025)
ArVoice: A Multi-Speaker Dataset for Arabic Speech Synthesis
von: Toyin, Hawau Olamide, et al.
Veröffentlicht: (2025)
von: Toyin, Hawau Olamide, et al.
Veröffentlicht: (2025)
SmoothSinger: A Conditional Diffusion Model for Singing Voice Synthesis with Multi-Resolution Architecture
von: Sui, Kehan, et al.
Veröffentlicht: (2025)
von: Sui, Kehan, et al.
Veröffentlicht: (2025)
A Real-Time Voice Activity Detection Based On Lightweight Neural
von: Jia, Jidong, et al.
Veröffentlicht: (2024)
von: Jia, Jidong, et al.
Veröffentlicht: (2024)
Lightweight Self-Supervised Detection of Fundamental Frequency and Accurate Probability of Voicing in Monophonic Music
von: Bitra, Venkat Suprabath, et al.
Veröffentlicht: (2026)
von: Bitra, Venkat Suprabath, et al.
Veröffentlicht: (2026)
TokenSynth: A Token-based Neural Synthesizer for Instrument Cloning and Text-to-Instrument
von: Kim, Kyungsu, et al.
Veröffentlicht: (2025)
von: Kim, Kyungsu, et al.
Veröffentlicht: (2025)
Interpreting Pretrained Speech Models for Automatic Speech Assessment of Voice Disorders
von: Lau, Hok-Shing, et al.
Veröffentlicht: (2024)
von: Lau, Hok-Shing, et al.
Veröffentlicht: (2024)
Vevo2: A Unified and Controllable Framework for Speech and Singing Voice Generation
von: Zhang, Xueyao, et al.
Veröffentlicht: (2025)
von: Zhang, Xueyao, et al.
Veröffentlicht: (2025)
LHQ-SVC: Lightweight and High Quality Singing Voice Conversion Modeling
von: Huang, Yubo, et al.
Veröffentlicht: (2024)
von: Huang, Yubo, et al.
Veröffentlicht: (2024)
YingMusic-Singer: Zero-shot Singing Voice Synthesis and Editing with Annotation-free Melody Guidance
von: Zheng, Junjie, et al.
Veröffentlicht: (2025)
von: Zheng, Junjie, et al.
Veröffentlicht: (2025)
Flowing Straighter with Conditional Flow Matching for Accurate Speech Enhancement
von: Cross, Mattias, et al.
Veröffentlicht: (2025)
von: Cross, Mattias, et al.
Veröffentlicht: (2025)
Look Once to Hear: Target Speech Hearing with Noisy Examples
von: Veluri, Bandhav, et al.
Veröffentlicht: (2024)
von: Veluri, Bandhav, et al.
Veröffentlicht: (2024)
i-LAVA: Insights on Low Latency Voice-2-Voice Architecture for Agents
von: Purwar, Anupam, et al.
Veröffentlicht: (2025)
von: Purwar, Anupam, et al.
Veröffentlicht: (2025)
IntrinsicVoice: Empowering LLMs with Intrinsic Real-time Voice Interaction Abilities
von: Zhang, Xin, et al.
Veröffentlicht: (2024)
von: Zhang, Xin, et al.
Veröffentlicht: (2024)
$τ$-Voice: Benchmarking Full-Duplex Voice Agents on Real-World Domains
von: Ray, Soham, et al.
Veröffentlicht: (2026)
von: Ray, Soham, et al.
Veröffentlicht: (2026)
CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models
von: Du, Zhihao, et al.
Veröffentlicht: (2024)
von: Du, Zhihao, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Voice Cloning for Dysarthric Speech Synthesis: Addressing Data Scarcity in Speech-Language Pathology
von: Moell, Birger, et al.
Veröffentlicht: (2025) -
Pronunciation Deviation Analysis Through Voice Cloning and Acoustic Comparison
von: Valdivia, Andrew, et al.
Veröffentlicht: (2025) -
SyncSpeech: Efficient and Low-Latency Text-to-Speech based on Temporal Masked Transformer
von: Sheng, Zhengyan, et al.
Veröffentlicht: (2025) -
Voice Cloning: Comprehensive Survey
von: Azzuni, Hussam, et al.
Veröffentlicht: (2025) -
Neural Multi-Speaker Voice Cloning for Nepali in Low-Resource Settings
von: Shrestha, Aayush M., et al.
Veröffentlicht: (2026)