Pashto Common Voice: Building the First Open Speech Corpus for a 60-Million-Speaker Low-Resource Language
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Rahman, Hanif, Rehman, Shafeeq ur |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
PashtoCorp: A 1.25-Billion-Word Corpus, Evaluation Suite, and Reproducible Pipeline for Low-Resource Language Development
von: Rahman, Hanif
Veröffentlicht: (2026)
von: Rahman, Hanif
Veröffentlicht: (2026)
Benchmarking Multilingual Speech Models on Pashto: Zero-Shot ASR, Script Failure, and Cross-Domain Evaluation
von: Rahman, Hanif
Veröffentlicht: (2026)
von: Rahman, Hanif
Veröffentlicht: (2026)
Fine-tuning Whisper for Pashto ASR: strategies and scale
von: Rahman, Hanif
Veröffentlicht: (2026)
von: Rahman, Hanif
Veröffentlicht: (2026)
PashtoTTS-Bench: automated screening for low-resource non-Latin-script text-to-speech
von: Rahman, Hanif
Veröffentlicht: (2026)
von: Rahman, Hanif
Veröffentlicht: (2026)
From Scarcity to Scale: A Release-Level Analysis of the Pashto Common Voice Dataset
von: Jahani, Jandad, et al.
Veröffentlicht: (2026)
von: Jahani, Jandad, et al.
Veröffentlicht: (2026)
Connecting Voices: LoReSpeech as a Low-Resource Speech Parallel Corpus
von: Ouzerrout, Samy
Veröffentlicht: (2025)
von: Ouzerrout, Samy
Veröffentlicht: (2025)
Saar-Voice: A Multi-Speaker Saarbrücken Dialect Speech Corpus
von: Oberkircher, Lena S., et al.
Veröffentlicht: (2026)
von: Oberkircher, Lena S., et al.
Veröffentlicht: (2026)
Towards Open Foundation Language Model and Corpus for Macedonian: A Low-Resource Language
von: Krsteski, Stefan, et al.
Veröffentlicht: (2025)
von: Krsteski, Stefan, et al.
Veröffentlicht: (2025)
Giving Voice to the Constitution: Low-Resource Text-to-Speech for Quechua and Spanish Using a Bilingual Legal Corpus
von: Ortega, John E., et al.
Veröffentlicht: (2026)
von: Ortega, John E., et al.
Veröffentlicht: (2026)
Developing an Open Conversational Speech Corpus for the Isan Language
von: Na-Thalang, Adisai, et al.
Veröffentlicht: (2025)
von: Na-Thalang, Adisai, et al.
Veröffentlicht: (2025)
SloPal: A 60-Million-Word Slovak Parliamentary Corpus with Aligned Speech and Fine-Tuned ASR Models
von: Božík, Erik, et al.
Veröffentlicht: (2025)
von: Božík, Erik, et al.
Veröffentlicht: (2025)
Building Efficient and Effective OpenQA Systems for Low-Resource Languages
von: Budur, Emrah, et al.
Veröffentlicht: (2024)
von: Budur, Emrah, et al.
Veröffentlicht: (2024)
SITA: Learning Speaker-Invariant and Tone-Aware Speech Representations for Low-Resource Tonal Languages
von: Xu, Tianyi, et al.
Veröffentlicht: (2026)
von: Xu, Tianyi, et al.
Veröffentlicht: (2026)
OpenWHO: A Document-Level Parallel Corpus for Health Translation in Low-Resource Languages
von: Merx, Raphaël, et al.
Veröffentlicht: (2025)
von: Merx, Raphaël, et al.
Veröffentlicht: (2025)
Quechua Speech Datasets in Common Voice: The Case of Puno Quechua
von: Huaman, Elwin, et al.
Veröffentlicht: (2025)
von: Huaman, Elwin, et al.
Veröffentlicht: (2025)
GlotCC: An Open Broad-Coverage CommonCrawl Corpus and Pipeline for Minority Languages
von: Kargaran, Amir Hossein, et al.
Veröffentlicht: (2024)
von: Kargaran, Amir Hossein, et al.
Veröffentlicht: (2024)
Speech-to-Speech Translation Pipelines for Conversations in Low-Resource Languages
von: Popescu-Belis, Andrei, et al.
Veröffentlicht: (2025)
von: Popescu-Belis, Andrei, et al.
Veröffentlicht: (2025)
Towards Inclusive ASR: Investigating Voice Conversion for Dysarthric Speech Recognition in Low-Resource Languages
von: Li, Chin-Jou, et al.
Veröffentlicht: (2025)
von: Li, Chin-Jou, et al.
Veröffentlicht: (2025)
MaiBERT: A Pre-training Corpus and Language Model for Low-Resourced Maithili Language
von: Yadav, Sumit, et al.
Veröffentlicht: (2025)
von: Yadav, Sumit, et al.
Veröffentlicht: (2025)
CorpusQA: A 10 Million Token Benchmark for Corpus-Level Analysis and Reasoning
von: Lu, Zhiyuan, et al.
Veröffentlicht: (2026)
von: Lu, Zhiyuan, et al.
Veröffentlicht: (2026)
Analysis of Speech Temporal Dynamics in the Context of Speaker Verification and Voice Anonymization
von: Tomashenko, Natalia, et al.
Veröffentlicht: (2024)
von: Tomashenko, Natalia, et al.
Veröffentlicht: (2024)
Enhancing Voice Wake-Up for Dysarthria: Mandarin Dysarthria Speech Corpus Release and Customized System Design
von: Gao, Ming, et al.
Veröffentlicht: (2024)
von: Gao, Ming, et al.
Veröffentlicht: (2024)
Low-Resource Safety Failures Are Action Failures, Not Representation Failures
von: Aziz, Rashad, et al.
Veröffentlicht: (2026)
von: Aziz, Rashad, et al.
Veröffentlicht: (2026)
DATASHI: A Parallel English-Tashlhiyt Corpus for Orthography Normalization and Low-Resource Language Processing
von: Monir, Nasser-Eddine, et al.
Veröffentlicht: (2026)
von: Monir, Nasser-Eddine, et al.
Veröffentlicht: (2026)
Overcoming Low-Resource Barriers in Tulu: Neural Models and Corpus Creation for OffensiveLanguage Identification
von: D, Anusha M, et al.
Veröffentlicht: (2025)
von: D, Anusha M, et al.
Veröffentlicht: (2025)
FalAR: A Large-scale Speaker-Annotated European Portuguese Speech Corpus of Parliamentary Sessions
von: Teixeira, Francisco, et al.
Veröffentlicht: (2026)
von: Teixeira, Francisco, et al.
Veröffentlicht: (2026)
Commonality and Individuality! Integrating Humor Commonality with Speaker Individuality for Humor Recognition
von: Zhu, Haohao, et al.
Veröffentlicht: (2025)
von: Zhu, Haohao, et al.
Veröffentlicht: (2025)
Speak & Improve Corpus 2025: an L2 English Speech Corpus for Language Assessment and Feedback
von: Knill, Kate, et al.
Veröffentlicht: (2024)
von: Knill, Kate, et al.
Veröffentlicht: (2024)
CV-18 NER: Augmented Common Voice for Named Entity Recognition from Arabic Speech
von: Saidi, Youssef, et al.
Veröffentlicht: (2026)
von: Saidi, Youssef, et al.
Veröffentlicht: (2026)
Building a Non-native Speech Corpus Featuring Chinese-English Bilingual Children: Compilation and Rationale
von: Hung, Hiuchung, et al.
Veröffentlicht: (2023)
von: Hung, Hiuchung, et al.
Veröffentlicht: (2023)
Assessing the Feasibility of Lightweight Whisper Models for Low-Resource Urdu Transcription
von: Antall, Abdul Rehman, et al.
Veröffentlicht: (2025)
von: Antall, Abdul Rehman, et al.
Veröffentlicht: (2025)
ArVoice: A Multi-Speaker Dataset for Arabic Speech Synthesis
von: Toyin, Hawau Olamide, et al.
Veröffentlicht: (2025)
von: Toyin, Hawau Olamide, et al.
Veröffentlicht: (2025)
BnTTS: Few-Shot Speaker Adaptation in Low-Resource Setting
von: Basher, Mohammad Jahid Ibna, et al.
Veröffentlicht: (2025)
von: Basher, Mohammad Jahid Ibna, et al.
Veröffentlicht: (2025)
The Thiomi Dataset: A Large-Scale Multimodal Corpus for Low-Resource African Languages
von: Mutisya, Hillary, et al.
Veröffentlicht: (2026)
von: Mutisya, Hillary, et al.
Veröffentlicht: (2026)
Mangosteen: An Open Thai Corpus for Language Model Pretraining
von: Phatthiyaphaibun, Wannaphong, et al.
Veröffentlicht: (2025)
von: Phatthiyaphaibun, Wannaphong, et al.
Veröffentlicht: (2025)
GigaSpeech 2: An Evolving, Large-Scale and Multi-domain ASR Corpus for Low-Resource Languages with Automated Crawling, Transcription and Refinement
von: Yang, Yifan, et al.
Veröffentlicht: (2024)
von: Yang, Yifan, et al.
Veröffentlicht: (2024)
FeruzaSpeech: A 60 Hour Uzbek Read Speech Corpus with Punctuation, Casing, and Context
von: Povey, Anna, et al.
Veröffentlicht: (2024)
von: Povey, Anna, et al.
Veröffentlicht: (2024)
Quantifying Geospatial in the Common Crawl Corpus
von: Ilyankou, Ilya, et al.
Veröffentlicht: (2024)
von: Ilyankou, Ilya, et al.
Veröffentlicht: (2024)
Small Models, Big Impact: Efficient Corpus and Graph-Based Adaptation of Small Multilingual Language Models for Low-Resource Languages
von: Gurgurov, Daniil, et al.
Veröffentlicht: (2025)
von: Gurgurov, Daniil, et al.
Veröffentlicht: (2025)
Adapting Multilingual LLMs to Low-Resource Languages using Continued Pre-training and Synthetic Corpus
von: Joshi, Raviraj, et al.
Veröffentlicht: (2024)
von: Joshi, Raviraj, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
PashtoCorp: A 1.25-Billion-Word Corpus, Evaluation Suite, and Reproducible Pipeline for Low-Resource Language Development
von: Rahman, Hanif
Veröffentlicht: (2026) -
Benchmarking Multilingual Speech Models on Pashto: Zero-Shot ASR, Script Failure, and Cross-Domain Evaluation
von: Rahman, Hanif
Veröffentlicht: (2026) -
Fine-tuning Whisper for Pashto ASR: strategies and scale
von: Rahman, Hanif
Veröffentlicht: (2026) -
PashtoTTS-Bench: automated screening for low-resource non-Latin-script text-to-speech
von: Rahman, Hanif
Veröffentlicht: (2026) -
From Scarcity to Scale: A Release-Level Analysis of the Pashto Common Voice Dataset
von: Jahani, Jandad, et al.
Veröffentlicht: (2026)