MultiGen: Child-Friendly Multilingual Speech Generator with LLMs

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Gao, Xiaoxue, Zhang, Huayun, Chen, Nancy F.
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914021484003328
author Gao, Xiaoxue
Zhang, Huayun
Chen, Nancy F.
author_facet Gao, Xiaoxue
Zhang, Huayun
Chen, Nancy F.
contents Generative speech models have demonstrated significant potential in improving human-machine interactions, offering valuable real-world applications such as language learning for children. However, achieving high-quality, child-friendly speech generation remains challenging, particularly for low-resource languages across diverse languages and cultural contexts. In this paper, we propose MultiGen, a multilingual speech generation model with child-friendly interaction, leveraging LLM architecture for speech generation tailored for low-resource languages. We propose to integrate age-appropriate multilingual speech generation using LLM architectures, which can be used to facilitate young children's communication with AI systems through culturally relevant context in three low-resource languages: Singaporean accent Mandarin, Malay, and Tamil. Experimental results from both objective metrics and subjective evaluations demonstrate the superior performance of the proposed MultiGen compared to baseline methods.
format Preprint
id arxiv_https___arxiv_org_abs_2508_08715
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MultiGen: Child-Friendly Multilingual Speech Generator with LLMs
Gao, Xiaoxue
Zhang, Huayun
Chen, Nancy F.
Audio and Speech Processing
Artificial Intelligence
Computation and Language
Signal Processing
Generative speech models have demonstrated significant potential in improving human-machine interactions, offering valuable real-world applications such as language learning for children. However, achieving high-quality, child-friendly speech generation remains challenging, particularly for low-resource languages across diverse languages and cultural contexts. In this paper, we propose MultiGen, a multilingual speech generation model with child-friendly interaction, leveraging LLM architecture for speech generation tailored for low-resource languages. We propose to integrate age-appropriate multilingual speech generation using LLM architectures, which can be used to facilitate young children's communication with AI systems through culturally relevant context in three low-resource languages: Singaporean accent Mandarin, Malay, and Tamil. Experimental results from both objective metrics and subjective evaluations demonstrate the superior performance of the proposed MultiGen compared to baseline methods.
title MultiGen: Child-Friendly Multilingual Speech Generator with LLMs
topic Audio and Speech Processing
Artificial Intelligence
Computation and Language
Signal Processing
url https://arxiv.org/abs/2508.08715