Developing multilingual speech synthesis system for Ojibwe, Mi'kmaq, and Maliseet
Fuente:
arXiv
Guardado en:
| Autores principales: | Wang, Shenran, Yang, Changbing, Parkhill, Mike, Quinn, Chad, Hammerly, Christopher, Zhu, Jian |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
The taste of IPA: Towards open-vocabulary keyword spotting and forced alignment in any language
por: Zhu, Jian, et al.
Publicado: (2023)
por: Zhu, Jian, et al.
Publicado: (2023)
Advancing LLM-based phoneme-to-grapheme for multilingual speech recognition
por: Dong, Lukuang, et al.
Publicado: (2026)
por: Dong, Lukuang, et al.
Publicado: (2026)
ZIPA: A family of efficient models for multilingual phone recognition
por: Zhu, Jian, et al.
Publicado: (2025)
por: Zhu, Jian, et al.
Publicado: (2025)
A dual task learning approach to fine-tune a multilingual semantic speech encoder for Spoken Language Understanding
por: Laperrière, Gaëlle, et al.
Publicado: (2024)
por: Laperrière, Gaëlle, et al.
Publicado: (2024)
emg2speech: Synthesizing speech from electromyography using self-supervised speech models
por: Gowda, Harshavardhana T., et al.
Publicado: (2025)
por: Gowda, Harshavardhana T., et al.
Publicado: (2025)
Improving child speech recognition with augmented child-like speech
por: Zhang, Yuanyuan, et al.
Publicado: (2024)
por: Zhang, Yuanyuan, et al.
Publicado: (2024)
Expressive paragraph text-to-speech synthesis with multi-step variational autoencoder
por: Li, Xuyuan, et al.
Publicado: (2023)
por: Li, Xuyuan, et al.
Publicado: (2023)
Translating speech with just images
por: Oneata, Dan, et al.
Publicado: (2024)
por: Oneata, Dan, et al.
Publicado: (2024)
Noise-robust zero-shot text-to-speech synthesis conditioned on self-supervised speech-representation model with adapters
por: Fujita, Kenichi, et al.
Publicado: (2024)
por: Fujita, Kenichi, et al.
Publicado: (2024)
The FruitShell French synthesis system at the Blizzard 2023 Challenge
por: Qi, Xin, et al.
Publicado: (2023)
por: Qi, Xin, et al.
Publicado: (2023)
Can large audio language models understand child stuttering speech? speech summarization, and source separation
por: Okocha, Chibuzor, et al.
Publicado: (2025)
por: Okocha, Chibuzor, et al.
Publicado: (2025)
A two-stage transliteration approach to improve performance of a multilingual ASR
por: Kumar, Rohit
Publicado: (2024)
por: Kumar, Rohit
Publicado: (2024)
A unified front-end framework for English text-to-speech synthesis
por: Ying, Zelin, et al.
Publicado: (2023)
por: Ying, Zelin, et al.
Publicado: (2023)
SpeechCLIP+: Self-supervised multi-task representation learning for speech via CLIP and speech-image data
por: Wang, Hsuan-Fu, et al.
Publicado: (2024)
por: Wang, Hsuan-Fu, et al.
Publicado: (2024)
Automated speech audiometry: Can it work using open-source pre-trained Kaldi-NL automatic speech recognition?
por: Araiza-Illan, Gloria, et al.
Publicado: (2023)
por: Araiza-Illan, Gloria, et al.
Publicado: (2023)
Semantic enrichment towards efficient speech representations
por: Laperrière, Gaëlle, et al.
Publicado: (2023)
por: Laperrière, Gaëlle, et al.
Publicado: (2023)
MiMo-Audio: Audio Language Models are Few-Shot Learners
por: Core Team, et al.
Publicado: (2025)
por: Core Team, et al.
Publicado: (2025)
Can Whisper perform speech-based in-context learning?
por: Wang, Siyin, et al.
Publicado: (2023)
por: Wang, Siyin, et al.
Publicado: (2023)
Revisiting speech segmentation and lexicon learning with better features
por: Kamper, Herman, et al.
Publicado: (2024)
por: Kamper, Herman, et al.
Publicado: (2024)
Bridging the gap: A comparative exploration of Speech-LLM and end-to-end architecture for multilingual conversational ASR
por: Mei, Yuxiang, et al.
Publicado: (2026)
por: Mei, Yuxiang, et al.
Publicado: (2026)
Self-consistent context aware conformer transducer for speech recognition
por: Kolokolov, Konstantin, et al.
Publicado: (2024)
por: Kolokolov, Konstantin, et al.
Publicado: (2024)
LLM-based phoneme-to-grapheme for phoneme-based speech recognition
por: Ma, Te, et al.
Publicado: (2025)
por: Ma, Te, et al.
Publicado: (2025)
An efficient text augmentation approach for contextualized Mandarin speech recognition
por: Zheng, Naijun, et al.
Publicado: (2024)
por: Zheng, Naijun, et al.
Publicado: (2024)
Direct Punjabi to English speech translation using discrete units
por: Kaur, Prabhjot, et al.
Publicado: (2024)
por: Kaur, Prabhjot, et al.
Publicado: (2024)
Transferable speech-to-text large language model alignment module
por: Wu, Boyong, et al.
Publicado: (2024)
por: Wu, Boyong, et al.
Publicado: (2024)
Transferring speech-generic and depression-specific knowledge for Alzheimer's disease detection
por: Cui, Ziyun, et al.
Publicado: (2023)
por: Cui, Ziyun, et al.
Publicado: (2023)
Natural language guidance of high-fidelity text-to-speech with synthetic annotations
por: Lyth, Dan, et al.
Publicado: (2024)
por: Lyth, Dan, et al.
Publicado: (2024)
Beyond the binary: Limitations and possibilities of gender-related speech technology research
por: Sanchez, Ariadna, et al.
Publicado: (2024)
por: Sanchez, Ariadna, et al.
Publicado: (2024)
An experiment on an automated literature survey of data-driven speech enhancement methods
por: Santos, Arthur dos, et al.
Publicado: (2023)
por: Santos, Arthur dos, et al.
Publicado: (2023)
Convoifilter: A case study of doing cocktail party speech recognition
por: Nguyen, Thai-Binh, et al.
Publicado: (2023)
por: Nguyen, Thai-Binh, et al.
Publicado: (2023)
Unsupervised lexicon learning from speech is limited by representations rather than clustering
por: Slabbert, Danel, et al.
Publicado: (2025)
por: Slabbert, Danel, et al.
Publicado: (2025)
Exploring the limits of decoder-only models trained on public speech recognition corpora
por: Gupta, Ankit, et al.
Publicado: (2024)
por: Gupta, Ankit, et al.
Publicado: (2024)
asr_eval: Algorithms and tools for multi-reference and streaming speech recognition evaluation
por: Sedukhin, Oleg, et al.
Publicado: (2026)
por: Sedukhin, Oleg, et al.
Publicado: (2026)
Data-driven grapheme-to-phoneme representations for a lexicon-free text-to-speech
por: Garg, Abhinav, et al.
Publicado: (2024)
por: Garg, Abhinav, et al.
Publicado: (2024)
Seed LiveInterpret 2.0: End-to-end Simultaneous Speech-to-speech Translation with Your Voice
por: Cheng, Shanbo, et al.
Publicado: (2025)
por: Cheng, Shanbo, et al.
Publicado: (2025)
Automatic speech recognition for the Nepali language using CNN, bidirectional LSTM and ResNet
por: Dhakal, Manish, et al.
Publicado: (2024)
por: Dhakal, Manish, et al.
Publicado: (2024)
Korean aegyo speech shows systematic F1 increase to signal childlike qualities
por: Kim, Ji-eun, et al.
Publicado: (2026)
por: Kim, Ji-eun, et al.
Publicado: (2026)
AfriHuBERT: A self-supervised speech representation model for African languages
por: Alabi, Jesujoba O., et al.
Publicado: (2024)
por: Alabi, Jesujoba O., et al.
Publicado: (2024)
Neighbors and relatives: How do speech embeddings reflect linguistic connections across the world?
por: Törö, Tuukka, et al.
Publicado: (2025)
por: Törö, Tuukka, et al.
Publicado: (2025)
Can we reconstruct a dysarthric voice with the large speech model Parler TTS?
por: Sanchez, Ariadna, et al.
Publicado: (2025)
por: Sanchez, Ariadna, et al.
Publicado: (2025)
Ejemplares similares
-
The taste of IPA: Towards open-vocabulary keyword spotting and forced alignment in any language
por: Zhu, Jian, et al.
Publicado: (2023) -
Advancing LLM-based phoneme-to-grapheme for multilingual speech recognition
por: Dong, Lukuang, et al.
Publicado: (2026) -
ZIPA: A family of efficient models for multilingual phone recognition
por: Zhu, Jian, et al.
Publicado: (2025) -
A dual task learning approach to fine-tune a multilingual semantic speech encoder for Spoken Language Understanding
por: Laperrière, Gaëlle, et al.
Publicado: (2024) -
emg2speech: Synthesizing speech from electromyography using self-supervised speech models
por: Gowda, Harshavardhana T., et al.
Publicado: (2025)