PashtoTTS-Bench: automated screening for low-resource non-Latin-script text-to-speech
Fuente:
arXiv
Guardado en:
| Autor principal: | Rahman, Hanif |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Strategies for improving low resource speech to text translation relying on pre-trained ASR models
por: Kesiraju, Santosh, et al.
Publicado: (2023)
por: Kesiraju, Santosh, et al.
Publicado: (2023)
Fine-tuning Whisper for Pashto ASR: strategies and scale
por: Rahman, Hanif
Publicado: (2026)
por: Rahman, Hanif
Publicado: (2026)
From Scarcity to Scale: A Release-Level Analysis of the Pashto Common Voice Dataset
por: Jahani, Jandad, et al.
Publicado: (2026)
por: Jahani, Jandad, et al.
Publicado: (2026)
Benchmarking Multilingual Speech Models on Pashto: Zero-Shot ASR, Script Failure, and Cross-Domain Evaluation
por: Rahman, Hanif
Publicado: (2026)
por: Rahman, Hanif
Publicado: (2026)
Can we reconstruct a dysarthric voice with the large speech model Parler TTS?
por: Sanchez, Ariadna, et al.
Publicado: (2025)
por: Sanchez, Ariadna, et al.
Publicado: (2025)
The Greek podcast corpus: Competitive speech models for low-resourced languages with weakly supervised data
por: Paraskevopoulos, Georgios, et al.
Publicado: (2024)
por: Paraskevopoulos, Georgios, et al.
Publicado: (2024)
An efficient text augmentation approach for contextualized Mandarin speech recognition
por: Zheng, Naijun, et al.
Publicado: (2024)
por: Zheng, Naijun, et al.
Publicado: (2024)
Transferable speech-to-text large language model alignment module
por: Wu, Boyong, et al.
Publicado: (2024)
por: Wu, Boyong, et al.
Publicado: (2024)
An experiment on an automated literature survey of data-driven speech enhancement methods
por: Santos, Arthur dos, et al.
Publicado: (2023)
por: Santos, Arthur dos, et al.
Publicado: (2023)
Natural language guidance of high-fidelity text-to-speech with synthetic annotations
por: Lyth, Dan, et al.
Publicado: (2024)
por: Lyth, Dan, et al.
Publicado: (2024)
PashtoCorp: A 1.25-Billion-Word Corpus, Evaluation Suite, and Reproducible Pipeline for Low-Resource Language Development
por: Rahman, Hanif
Publicado: (2026)
por: Rahman, Hanif
Publicado: (2026)
Data-driven grapheme-to-phoneme representations for a lexicon-free text-to-speech
por: Garg, Abhinav, et al.
Publicado: (2024)
por: Garg, Abhinav, et al.
Publicado: (2024)
EE-TTS: Emphatic Expressive TTS with Linguistic Information
por: Zhong, Yi, et al.
Publicado: (2023)
por: Zhong, Yi, et al.
Publicado: (2023)
Pashto Common Voice: Building the First Open Speech Corpus for a 60-Million-Speaker Low-Resource Language
por: Rahman, Hanif, et al.
Publicado: (2026)
por: Rahman, Hanif, et al.
Publicado: (2026)
MOSS-TTS Technical Report
por: Gong, Yitian, et al.
Publicado: (2026)
por: Gong, Yitian, et al.
Publicado: (2026)
Tibetan-TTS:Low-Resource Tibetan Speech Synthesis with Large Model Adaptation
por: He, Jiaxu, et al.
Publicado: (2026)
por: He, Jiaxu, et al.
Publicado: (2026)
Noise-robust zero-shot text-to-speech synthesis conditioned on self-supervised speech-representation model with adapters
por: Fujita, Kenichi, et al.
Publicado: (2024)
por: Fujita, Kenichi, et al.
Publicado: (2024)
TouchTTS: An Embarrassingly Simple TTS Framework that Everyone Can Touch
por: Song, Xingchen, et al.
Publicado: (2024)
por: Song, Xingchen, et al.
Publicado: (2024)
Transfer the linguistic representations from TTS to accent conversion with non-parallel data
por: Chen, Xi, et al.
Publicado: (2024)
por: Chen, Xi, et al.
Publicado: (2024)
RephraseTTS: Dynamic Length Text based Speech Insertion with Speaker Style Transfer
por: Matiyali, Neeraj, et al.
Publicado: (2025)
por: Matiyali, Neeraj, et al.
Publicado: (2025)
A unified front-end framework for English text-to-speech synthesis
por: Ying, Zelin, et al.
Publicado: (2023)
por: Ying, Zelin, et al.
Publicado: (2023)
Training dynamic models using early exits for automatic speech recognition on resource-constrained devices
por: Wright, George August, et al.
Publicado: (2023)
por: Wright, George August, et al.
Publicado: (2023)
Covertly improving intelligibility with data-driven adaptations of speech timing
por: Tuttösí, Paige, et al.
Publicado: (2026)
por: Tuttösí, Paige, et al.
Publicado: (2026)
Qwen3-TTS Technical Report
por: Hu, Hangrui, et al.
Publicado: (2026)
por: Hu, Hangrui, et al.
Publicado: (2026)
A2TTS: TTS for Low Resource Indian Languages
por: Bhadoriya, Ayush Singh, et al.
Publicado: (2025)
por: Bhadoriya, Ayush Singh, et al.
Publicado: (2025)
The TTS-STT Flywheel: Synthetic Entity-Dense Audio Closes the Indic ASR Gap Where Commercial and Open-Source Systems Fail
por: Menta, Venkata Pushpak Teja
Publicado: (2026)
por: Menta, Venkata Pushpak Teja
Publicado: (2026)
emg2speech: Synthesizing speech from electromyography using self-supervised speech models
por: Gowda, Harshavardhana T., et al.
Publicado: (2025)
por: Gowda, Harshavardhana T., et al.
Publicado: (2025)
Moshi: a speech-text foundation model for real-time dialogue
por: Défossez, Alexandre, et al.
Publicado: (2024)
por: Défossez, Alexandre, et al.
Publicado: (2024)
Improving child speech recognition with augmented child-like speech
por: Zhang, Yuanyuan, et al.
Publicado: (2024)
por: Zhang, Yuanyuan, et al.
Publicado: (2024)
Multi-interaction TTS toward professional recording reproduction
por: Kanagawa, Hiroki, et al.
Publicado: (2025)
por: Kanagawa, Hiroki, et al.
Publicado: (2025)
MunTTS: A Text-to-Speech System for Mundari
por: Gumma, Varun, et al.
Publicado: (2024)
por: Gumma, Varun, et al.
Publicado: (2024)
Llama-VITS: Enhancing TTS Synthesis with Semantic Awareness
por: Feng, Xincan, et al.
Publicado: (2024)
por: Feng, Xincan, et al.
Publicado: (2024)
RWKVTTS: Yet another TTS based on RWKV-7
por: yueyu, Lin, et al.
Publicado: (2025)
por: yueyu, Lin, et al.
Publicado: (2025)
DiaMoE-TTS: A Unified IPA-Based Dialect TTS Framework with Mixture-of-Experts and Parameter-Efficient Zero-Shot Adaptation
por: Chen, Ziqi, et al.
Publicado: (2025)
por: Chen, Ziqi, et al.
Publicado: (2025)
Low-resource speech recognition and dialect identification of Irish in a multi-task framework
por: Lonergan, Liam, et al.
Publicado: (2024)
por: Lonergan, Liam, et al.
Publicado: (2024)
Assessing the Ability of Neural TTS Systems to Model Consonant-Induced F0 Perturbation
por: Yang, Tianle, et al.
Publicado: (2026)
por: Yang, Tianle, et al.
Publicado: (2026)
Word-wise intonation model for cross-language TTS systems
por: A., Tomilov A., et al.
Publicado: (2024)
por: A., Tomilov A., et al.
Publicado: (2024)
An investigation of phrase break prediction in an End-to-End TTS system
por: Vadapalli, Anandaswarup
Publicado: (2023)
por: Vadapalli, Anandaswarup
Publicado: (2023)
JoyTTS: LLM-based Spoken Chatbot With Voice Cloning
por: Zhou, Fangru, et al.
Publicado: (2025)
por: Zhou, Fangru, et al.
Publicado: (2025)
A Language Modeling Approach to Diacritic-Free Hebrew TTS
por: Roth, Amit, et al.
Publicado: (2024)
por: Roth, Amit, et al.
Publicado: (2024)
Ejemplares similares
-
Strategies for improving low resource speech to text translation relying on pre-trained ASR models
por: Kesiraju, Santosh, et al.
Publicado: (2023) -
Fine-tuning Whisper for Pashto ASR: strategies and scale
por: Rahman, Hanif
Publicado: (2026) -
From Scarcity to Scale: A Release-Level Analysis of the Pashto Common Voice Dataset
por: Jahani, Jandad, et al.
Publicado: (2026) -
Benchmarking Multilingual Speech Models on Pashto: Zero-Shot ASR, Script Failure, and Cross-Domain Evaluation
por: Rahman, Hanif
Publicado: (2026) -
Can we reconstruct a dysarthric voice with the large speech model Parler TTS?
por: Sanchez, Ariadna, et al.
Publicado: (2025)