Neural Multi-Speaker Voice Cloning for Nepali in Low-Resource Settings

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shrestha, Aayush M., Bajracharya, Aditya, Shakya, Projan, Kshatri, Dinesh B.
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908788954497024
author Shrestha, Aayush M.
Bajracharya, Aditya
Shakya, Projan
Kshatri, Dinesh B.
author_facet Shrestha, Aayush M.
Bajracharya, Aditya
Shakya, Projan
Kshatri, Dinesh B.
contents This research presents a few-shot voice cloning system for Nepali speakers, designed to synthesize speech in a specific speaker's voice from Devanagari text using minimal data. Voice cloning in Nepali remains largely unexplored due to its low-resource nature. To address this, we constructed separate datasets: untranscribed audio for training a speaker encoder and paired text-audio data for training a Tacotron2-based synthesizer. The speaker encoder, optimized with Generative End2End loss, generates embeddings that capture the speaker's vocal identity, validated through Uniform Manifold Approximation and Projection (UMAP) for dimension reduction visualizations. These embeddings are fused with Tacotron2's text embeddings to produce mel-spectrograms, which are then converted into audio using a WaveRNN vocoder. Audio data were collected from various sources, including self-recordings, and underwent thorough preprocessing for quality and alignment. Training was performed using mel and gate loss functions under multiple hyperparameter settings. The system effectively clones speaker characteristics even for unseen voices, demonstrating the feasibility of few-shot voice cloning for the Nepali language and establishing a foundation for personalized speech synthesis in low-resource scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2601_18694
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Neural Multi-Speaker Voice Cloning for Nepali in Low-Resource Settings
Shrestha, Aayush M.
Bajracharya, Aditya
Shakya, Projan
Kshatri, Dinesh B.
Sound
Artificial Intelligence
This research presents a few-shot voice cloning system for Nepali speakers, designed to synthesize speech in a specific speaker's voice from Devanagari text using minimal data. Voice cloning in Nepali remains largely unexplored due to its low-resource nature. To address this, we constructed separate datasets: untranscribed audio for training a speaker encoder and paired text-audio data for training a Tacotron2-based synthesizer. The speaker encoder, optimized with Generative End2End loss, generates embeddings that capture the speaker's vocal identity, validated through Uniform Manifold Approximation and Projection (UMAP) for dimension reduction visualizations. These embeddings are fused with Tacotron2's text embeddings to produce mel-spectrograms, which are then converted into audio using a WaveRNN vocoder. Audio data were collected from various sources, including self-recordings, and underwent thorough preprocessing for quality and alignment. Training was performed using mel and gate loss functions under multiple hyperparameter settings. The system effectively clones speaker characteristics even for unseen voices, demonstrating the feasibility of few-shot voice cloning for the Nepali language and establishing a foundation for personalized speech synthesis in low-resource scenarios.
title Neural Multi-Speaker Voice Cloning for Nepali in Low-Resource Settings
topic Sound
Artificial Intelligence
url https://arxiv.org/abs/2601.18694