MultiGen: Child-Friendly Multilingual Speech Generator with LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Gao, Xiaoxue, Zhang, Huayun, Chen, Nancy F. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Prompt-Unseen-Emotion: Zero-shot Expressive Speech Synthesis with Prompt-LLM Contextual Knowledge for Mixed Emotions
by: Gao, Xiaoxue, et al.
Published: (2025)
by: Gao, Xiaoxue, et al.
Published: (2025)
Emo-DPO: Controllable Emotional Speech Synthesis through Direct Preference Optimization
by: Gao, Xiaoxue, et al.
Published: (2024)
by: Gao, Xiaoxue, et al.
Published: (2024)
Crowdsourced Multilingual Speech Intelligibility Testing
by: Lechler, Laura, et al.
Published: (2024)
by: Lechler, Laura, et al.
Published: (2024)
Improved Cross-Lingual Transfer Learning For Automatic Speech Translation
by: Khurana, Sameer, et al.
Published: (2023)
by: Khurana, Sameer, et al.
Published: (2023)
Enhancing Listened Speech Decoding from EEG via Parallel Phoneme Sequence Prediction
by: Lee, Jihwan, et al.
Published: (2025)
by: Lee, Jihwan, et al.
Published: (2025)
Textless Unit-to-Unit training for Many-to-Many Multilingual Speech-to-Speech Translation
by: Kim, Minsu, et al.
Published: (2023)
by: Kim, Minsu, et al.
Published: (2023)
TTSlow: Slow Down Text-to-Speech with Efficiency Robustness Evaluations
by: Gao, Xiaoxue, et al.
Published: (2024)
by: Gao, Xiaoxue, et al.
Published: (2024)
Transferable Adversarial Attacks against ASR
by: Gao, Xiaoxue, et al.
Published: (2024)
by: Gao, Xiaoxue, et al.
Published: (2024)
Reading Miscue Detection in Primary School through Automatic Speech Recognition
by: Gao, Lingyun, et al.
Published: (2024)
by: Gao, Lingyun, et al.
Published: (2024)
Automatic Speech Recognition of Non-Native Child Speech for Language Learning Applications
by: Wills, Simone, et al.
Published: (2023)
by: Wills, Simone, et al.
Published: (2023)
Confidence-Based Self-Training for EMG-to-Speech: Leveraging Synthetic EMG for Robust Modeling
by: Chen, Xiaodan, et al.
Published: (2025)
by: Chen, Xiaodan, et al.
Published: (2025)
Tone recognition in low-resource languages of North-East India: peeling the layers of SSL-based speech models
by: Gogoi, Parismita, et al.
Published: (2025)
by: Gogoi, Parismita, et al.
Published: (2025)
Is one brick enough to break the wall of spoken dialogue state tracking?
by: Druart, Lucas, et al.
Published: (2023)
by: Druart, Lucas, et al.
Published: (2023)
Language Bias in Self-Supervised Learning For Automatic Speech Recognition
by: Storey, Edward, et al.
Published: (2025)
by: Storey, Edward, et al.
Published: (2025)
GenTranslate: Large Language Models are Generative Multilingual Speech and Machine Translators
by: Hu, Yuchen, et al.
Published: (2024)
by: Hu, Yuchen, et al.
Published: (2024)
Utterance-Level Methods for Identifying Reliable ASR-Output for Child Speech
by: Lathouwers, Gus, et al.
Published: (2026)
by: Lathouwers, Gus, et al.
Published: (2026)
TRI-DEP: A Trimodal Comparative Study for Depression Detection Using Speech, Text, and EEG
by: Nurfidausi, Annisaa Fitri, et al.
Published: (2025)
by: Nurfidausi, Annisaa Fitri, et al.
Published: (2025)
CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation
by: Kim, Ji-Hoon, et al.
Published: (2024)
by: Kim, Ji-Hoon, et al.
Published: (2024)
Mitigating Intra-Speaker Variability in Diarization with Style-Controllable Speech Augmentation
by: Kim, Miseul, et al.
Published: (2025)
by: Kim, Miseul, et al.
Published: (2025)
TouchASP: Elastic Automatic Speech Perception that Everyone Can Touch
by: Song, Xingchen, et al.
Published: (2024)
by: Song, Xingchen, et al.
Published: (2024)
Advancing Speech Summarization in Multi-modal LLMs with Reinforcement Learning
by: Ling, Shaoshi, et al.
Published: (2025)
by: Ling, Shaoshi, et al.
Published: (2025)
MORE: Multi-Objective Adversarial Attacks on Speech Recognition
by: Gao, Xiaoxue, et al.
Published: (2026)
by: Gao, Xiaoxue, et al.
Published: (2026)
WST-X Series: Wavelet Scattering Transform for Interpretable Speech Deepfake Detection
by: Xuan, Xi, et al.
Published: (2026)
by: Xuan, Xi, et al.
Published: (2026)
Visual-Aware Speech Recognition for Noisy Scenarios
by: Balaji, Lakshmipathi, et al.
Published: (2025)
by: Balaji, Lakshmipathi, et al.
Published: (2025)
Privacy-Preserving End-to-End Full-Duplex Speech Dialogue Models
by: Kuzmin, Nikita, et al.
Published: (2026)
by: Kuzmin, Nikita, et al.
Published: (2026)
WaveSP-Net: Learnable Wavelet-Domain Sparse Prompt Tuning for Speech Deepfake Detection
by: Xuan, Xi, et al.
Published: (2025)
by: Xuan, Xi, et al.
Published: (2025)
A Large-Scale Evaluation of Speech Foundation Models
by: Yang, Shu-wen, et al.
Published: (2024)
by: Yang, Shu-wen, et al.
Published: (2024)
Transsion Multilingual Speech Recognition System for MLC-SLM 2025 Challenge
by: Li, Xiaoxiao, et al.
Published: (2025)
by: Li, Xiaoxiao, et al.
Published: (2025)
WAXAL: A Large-Scale Multilingual African Language Speech Corpus
by: Diack, Abdoulaye, et al.
Published: (2026)
by: Diack, Abdoulaye, et al.
Published: (2026)
Improving Child Speech Recognition and Reading Mistake Detection by Using Prompts
by: Gao, Lingyun, et al.
Published: (2025)
by: Gao, Lingyun, et al.
Published: (2025)
Voice Mapping of Text-to-Speech Systems: A Metric-Based Approach for Voice Quality Assessment
by: Cai, Huanchen, et al.
Published: (2026)
by: Cai, Huanchen, et al.
Published: (2026)
FLEURS-R: A Restored Multilingual Speech Corpus for Generation Tasks
by: Ma, Min, et al.
Published: (2024)
by: Ma, Min, et al.
Published: (2024)
Speech Boosting: Low-Latency Live Speech Enhancement for TWS Earbuds
by: Bae, Hanbin, et al.
Published: (2024)
by: Bae, Hanbin, et al.
Published: (2024)
EEG-Based Speech Decoding: A Novel Approach Using Multi-Kernel Ensemble Diffusion Models
by: Kim, Soowon, et al.
Published: (2024)
by: Kim, Soowon, et al.
Published: (2024)
Laugh Now Cry Later: Controlling Time-Varying Emotional States of Flow-Matching-Based Zero-Shot Text-to-Speech
by: Wu, Haibin, et al.
Published: (2024)
by: Wu, Haibin, et al.
Published: (2024)
EmoBox: Multilingual Multi-corpus Speech Emotion Recognition Toolkit and Benchmark
by: Ma, Ziyang, et al.
Published: (2024)
by: Ma, Ziyang, et al.
Published: (2024)
Construction and Evaluation of Mandarin Multimodal Emotional Speech Database
by: Ting, Zhu, et al.
Published: (2024)
by: Ting, Zhu, et al.
Published: (2024)
Can Speech LLMs Think while Listening?
by: Shih, Yi-Jen, et al.
Published: (2025)
by: Shih, Yi-Jen, et al.
Published: (2025)
A Unified Speech LLM for Diarization and Speech Recognition in Multilingual Conversations
by: Saengthong, Phurich, et al.
Published: (2025)
by: Saengthong, Phurich, et al.
Published: (2025)
SIFT-50M: A Large-Scale Multilingual Dataset for Speech Instruction Fine-Tuning
by: Pandey, Prabhat, et al.
Published: (2025)
by: Pandey, Prabhat, et al.
Published: (2025)
Similar Items
-
Prompt-Unseen-Emotion: Zero-shot Expressive Speech Synthesis with Prompt-LLM Contextual Knowledge for Mixed Emotions
by: Gao, Xiaoxue, et al.
Published: (2025) -
Emo-DPO: Controllable Emotional Speech Synthesis through Direct Preference Optimization
by: Gao, Xiaoxue, et al.
Published: (2024) -
Crowdsourced Multilingual Speech Intelligibility Testing
by: Lechler, Laura, et al.
Published: (2024) -
Improved Cross-Lingual Transfer Learning For Automatic Speech Translation
by: Khurana, Sameer, et al.
Published: (2023) -
Enhancing Listened Speech Decoding from EEG via Parallel Phoneme Sequence Prediction
by: Lee, Jihwan, et al.
Published: (2025)