Neural Multi-Speaker Voice Cloning for Nepali in Low-Resource Settings
Fuente:
arXiv
Saved in:
| Main Authors: | Shrestha, Aayush M., Bajracharya, Aditya, Shakya, Projan, Kshatri, Dinesh B. |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Advancing Voice Cloning for Nepali: Leveraging Transfer Learning in a Low-Resource Language
by: Karki, Manjil, et al.
Published: (2024)
by: Karki, Manjil, et al.
Published: (2024)
i-LAVA: Insights on Low Latency Voice-2-Voice Architecture for Agents
by: Purwar, Anupam, et al.
Published: (2025)
by: Purwar, Anupam, et al.
Published: (2025)
VoiceMark: Zero-Shot Voice Cloning-Resistant Watermarking Approach Leveraging Speaker-Specific Latents
by: Li, Haiyun, et al.
Published: (2025)
by: Li, Haiyun, et al.
Published: (2025)
Voice Cloning: Comprehensive Survey
by: Azzuni, Hussam, et al.
Published: (2025)
by: Azzuni, Hussam, et al.
Published: (2025)
Proactive Detection of Voice Cloning with Localized Watermarking
by: Roman, Robin San, et al.
Published: (2024)
by: Roman, Robin San, et al.
Published: (2024)
A Lightweight Pipeline for Noisy Speech Voice Cloning and Accurate Lip Sync Synthesis
by: Amir, Javeria, et al.
Published: (2025)
by: Amir, Javeria, et al.
Published: (2025)
The ISCSLP 2024 Conversational Voice Clone (CoVoC) Challenge: Tasks, Results and Findings
by: Xia, Kangxiang, et al.
Published: (2024)
by: Xia, Kangxiang, et al.
Published: (2024)
CloneShield: A Framework for Universal Perturbation Against Zero-Shot Voice Cloning
by: Li, Renyuan, et al.
Published: (2025)
by: Li, Renyuan, et al.
Published: (2025)
Voice "Cloning" is Style Transfer
by: Zhou, Kaitlyn, et al.
Published: (2026)
by: Zhou, Kaitlyn, et al.
Published: (2026)
Pronunciation Deviation Analysis Through Voice Cloning and Acoustic Comparison
by: Valdivia, Andrew, et al.
Published: (2025)
by: Valdivia, Andrew, et al.
Published: (2025)
Probabilistic Fusion and Calibration of Neural Speaker Diarization Models
by: Alvarez-Trejos, Juan Ignacio, et al.
Published: (2025)
by: Alvarez-Trejos, Juan Ignacio, et al.
Published: (2025)
TokenSynth: A Token-based Neural Synthesizer for Instrument Cloning and Text-to-Instrument
by: Kim, Kyungsu, et al.
Published: (2025)
by: Kim, Kyungsu, et al.
Published: (2025)
X-Voice: Enabling Everyone to Speak 30 Languages via Zero-Shot Cross-Lingual Voice Cloning
by: Xu, Rixi, et al.
Published: (2026)
by: Xu, Rixi, et al.
Published: (2026)
MultiActor-Audiobook: Zero-Shot Audiobook Generation with Faces and Voices of Multiple Speakers
by: Park, Kyeongman, et al.
Published: (2025)
by: Park, Kyeongman, et al.
Published: (2025)
ArVoice: A Multi-Speaker Dataset for Arabic Speech Synthesis
by: Toyin, Hawau Olamide, et al.
Published: (2025)
by: Toyin, Hawau Olamide, et al.
Published: (2025)
Fed-PISA: Federated Voice Cloning via Personalized Identity-Style Adaptation
by: Wang, Qi, et al.
Published: (2025)
by: Wang, Qi, et al.
Published: (2025)
Voice Cloning for Dysarthric Speech Synthesis: Addressing Data Scarcity in Speech-Language Pathology
by: Moell, Birger, et al.
Published: (2025)
by: Moell, Birger, et al.
Published: (2025)
Asynchronous Voice Anonymization Using Adversarial Perturbation On Speaker Embedding
by: Wang, Rui, et al.
Published: (2024)
by: Wang, Rui, et al.
Published: (2024)
NepaliGPT: A Generative Language Model for the Nepali Language
by: Pudasaini, Shushanta, et al.
Published: (2025)
by: Pudasaini, Shushanta, et al.
Published: (2025)
Self Voice Conversion as an Attack against Neural Audio Watermarking
by: Özer, Yigitcan, et al.
Published: (2026)
by: Özer, Yigitcan, et al.
Published: (2026)
Rethinking Leveraging Pre-Trained Multi-Layer Representations for Speaker Verification
by: Kim, Jin Sob, et al.
Published: (2025)
by: Kim, Jin Sob, et al.
Published: (2025)
Towards Nepali-language LLMs: Efficient GPT training with a Nepali BPE tokenizer
by: Shrestha, Adarsha, et al.
Published: (2025)
by: Shrestha, Adarsha, et al.
Published: (2025)
End-to-End Multi-Microphone Speaker Extraction Using Relative Transfer Functions
by: Eisenberg, Aviad, et al.
Published: (2025)
by: Eisenberg, Aviad, et al.
Published: (2025)
Robust Target Speaker Diarization and Separation via Augmented Speaker Embedding Sampling
by: Jalal, Md Asif, et al.
Published: (2025)
by: Jalal, Md Asif, et al.
Published: (2025)
VoiceCloak: A Multi-Dimensional Defense Framework against Unauthorized Diffusion-based Voice Cloning
by: Hu, Qianyue, et al.
Published: (2025)
by: Hu, Qianyue, et al.
Published: (2025)
SpeakerLM: End-to-End Versatile Speaker Diarization and Recognition with Multimodal Large Language Models
by: Yin, Han, et al.
Published: (2025)
by: Yin, Han, et al.
Published: (2025)
Do Not Mimic My Voice: Speaker Identity Unlearning for Zero-Shot Text-to-Speech
by: Kim, Taesoo, et al.
Published: (2025)
by: Kim, Taesoo, et al.
Published: (2025)
Leveraging Speaker Embeddings in End-to-End Neural Diarization for Two-Speaker Scenarios
by: Alvarez-Trejos, Juan Ignacio, et al.
Published: (2024)
by: Alvarez-Trejos, Juan Ignacio, et al.
Published: (2024)
Exploring Speaker Diarization with Mixture of Experts
by: Yang, Gaobin, et al.
Published: (2025)
by: Yang, Gaobin, et al.
Published: (2025)
MindVoice: Reconstructing Intelligible Speech from Non-invasive Neural Signals with Pretrained Priors
by: Bao, Guangyin, et al.
Published: (2026)
by: Bao, Guangyin, et al.
Published: (2026)
Disentangling Speakers in Multi-Talker Speech Recognition with Speaker-Aware CTC
by: Kang, Jiawen, et al.
Published: (2024)
by: Kang, Jiawen, et al.
Published: (2024)
HQ-SVC: Towards High-Quality Zero-Shot Singing Voice Conversion in Low-Resource Scenarios
by: Bai, Bingsong, et al.
Published: (2025)
by: Bai, Bingsong, et al.
Published: (2025)
$τ$-Voice: Benchmarking Full-Duplex Voice Agents on Real-World Domains
by: Ray, Soham, et al.
Published: (2026)
by: Ray, Soham, et al.
Published: (2026)
Continual Speaker Identity Unlearning with Minimal Interference
by: Kim, Jinju, et al.
Published: (2026)
by: Kim, Jinju, et al.
Published: (2026)
IntrinsicVoice: Empowering LLMs with Intrinsic Real-time Voice Interaction Abilities
by: Zhang, Xin, et al.
Published: (2024)
by: Zhang, Xin, et al.
Published: (2024)
LASPA: Language Agnostic Speaker Disentanglement with Prefix-Tuned Cross-Attention
by: Menon, Aditya Srinivas, et al.
Published: (2025)
by: Menon, Aditya Srinivas, et al.
Published: (2025)
SmoothSinger: A Conditional Diffusion Model for Singing Voice Synthesis with Multi-Resolution Architecture
by: Sui, Kehan, et al.
Published: (2025)
by: Sui, Kehan, et al.
Published: (2025)
Speaker Verification with Speech-Aware LLMs: Evaluation and Augmentation
by: Thebaud, Thomas, et al.
Published: (2026)
by: Thebaud, Thomas, et al.
Published: (2026)
Evaluating Identity Leakage in Speaker De-Identification Systems
by: Seo, Seungmin, et al.
Published: (2025)
by: Seo, Seungmin, et al.
Published: (2025)
Improving Low-Resource Dialect Classification Using Retrieval-based Voice Conversion
by: Fischbach, Lea, et al.
Published: (2025)
by: Fischbach, Lea, et al.
Published: (2025)
Similar Items
-
Advancing Voice Cloning for Nepali: Leveraging Transfer Learning in a Low-Resource Language
by: Karki, Manjil, et al.
Published: (2024) -
i-LAVA: Insights on Low Latency Voice-2-Voice Architecture for Agents
by: Purwar, Anupam, et al.
Published: (2025) -
VoiceMark: Zero-Shot Voice Cloning-Resistant Watermarking Approach Leveraging Speaker-Specific Latents
by: Li, Haiyun, et al.
Published: (2025) -
Voice Cloning: Comprehensive Survey
by: Azzuni, Hussam, et al.
Published: (2025) -
Proactive Detection of Voice Cloning with Localized Watermarking
by: Roman, Robin San, et al.
Published: (2024)