Developing an Open Conversational Speech Corpus for the Isan Language
Fuente:
arXiv
Saved in:
| Main Authors: | Na-Thalang, Adisai, Wittayasakpan, Chanakan, Phatcharoen, Kritsadha, Buakaw, Supakit |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ThaiOCRBench: A Task-Diverse Benchmark for Vision-Language Understanding in Thai
by: Nonesung, Surapon, et al.
Published: (2025)
by: Nonesung, Surapon, et al.
Published: (2025)
Typhoon ASR Real-time: FastConformer-Transducer for Thai Automatic Speech Recognition
by: Sirichotedumrong, Warit, et al.
Published: (2026)
by: Sirichotedumrong, Warit, et al.
Published: (2026)
Typhoon 2: A Family of Open Text and Multimodal Thai Large Language Models
by: Pipatanakul, Kunat, et al.
Published: (2024)
by: Pipatanakul, Kunat, et al.
Published: (2024)
Advancing Speech Translation: A Corpus of Mandarin-English Conversational Telephone Speech
by: Wotherspoon, Shannon, et al.
Published: (2024)
by: Wotherspoon, Shannon, et al.
Published: (2024)
Speak & Improve Corpus 2025: an L2 English Speech Corpus for Language Assessment and Feedback
by: Knill, Kate, et al.
Published: (2024)
by: Knill, Kate, et al.
Published: (2024)
Pashto Common Voice: Building the First Open Speech Corpus for a 60-Million-Speaker Low-Resource Language
by: Rahman, Hanif, et al.
Published: (2026)
by: Rahman, Hanif, et al.
Published: (2026)
Mangosteen: An Open Thai Corpus for Language Model Pretraining
by: Phatthiyaphaibun, Wannaphong, et al.
Published: (2025)
by: Phatthiyaphaibun, Wannaphong, et al.
Published: (2025)
Speech-to-Speech Translation Pipelines for Conversations in Low-Resource Languages
by: Popescu-Belis, Andrei, et al.
Published: (2025)
by: Popescu-Belis, Andrei, et al.
Published: (2025)
EuroSpeech: A Multilingual Speech Corpus
by: Pfisterer, Samuel, et al.
Published: (2025)
by: Pfisterer, Samuel, et al.
Published: (2025)
Towards Open Foundation Language Model and Corpus for Macedonian: A Low-Resource Language
by: Krsteski, Stefan, et al.
Published: (2025)
by: Krsteski, Stefan, et al.
Published: (2025)
FFSTC: Fongbe to French Speech Translation Corpus
by: Kponou, D. Fortune, et al.
Published: (2024)
by: Kponou, D. Fortune, et al.
Published: (2024)
Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research
by: Soldaini, Luca, et al.
Published: (2024)
by: Soldaini, Luca, et al.
Published: (2024)
Developing Conversational Speech Systems for Robots to Detect Speech Biomarkers of Cognition in People Living with Dementia
by: Perumandla, Rohith, et al.
Published: (2025)
by: Perumandla, Rohith, et al.
Published: (2025)
An Annotated Corpus of Arabic Tweets for Hate Speech Analysis
by: Zaghouani, Wajdi, et al.
Published: (2025)
by: Zaghouani, Wajdi, et al.
Published: (2025)
MzansiText and MzansiLM: An Open Corpus and Decoder-Only Language Model for South African Languages
by: Lombard, Anri, et al.
Published: (2026)
by: Lombard, Anri, et al.
Published: (2026)
RegSpeech12: A Regional Corpus of Bengali Spontaneous Speech Across Dialects
by: Hassan, Md. Rezuwan, et al.
Published: (2025)
by: Hassan, Md. Rezuwan, et al.
Published: (2025)
Connecting Voices: LoReSpeech as a Low-Resource Speech Parallel Corpus
by: Ouzerrout, Samy
Published: (2025)
by: Ouzerrout, Samy
Published: (2025)
RAGTruth: A Hallucination Corpus for Developing Trustworthy Retrieval-Augmented Language Models
by: Niu, Cheng, et al.
Published: (2023)
by: Niu, Cheng, et al.
Published: (2023)
ParCzech4Speech: A New Speech Corpus Derived from Czech Parliamentary Data
by: Stankov, Vladislav, et al.
Published: (2025)
by: Stankov, Vladislav, et al.
Published: (2025)
Saar-Voice: A Multi-Speaker Saarbrücken Dialect Speech Corpus
by: Oberkircher, Lena S., et al.
Published: (2026)
by: Oberkircher, Lena S., et al.
Published: (2026)
ZAEBUC-Spoken: A Multilingual Multidialectal Arabic-English Speech Corpus
by: Hamed, Injy, et al.
Published: (2024)
by: Hamed, Injy, et al.
Published: (2024)
ÌròyìnSpeech: A multi-purpose Yorùbá Speech Corpus
by: Ogunremi, Tolulope, et al.
Published: (2023)
by: Ogunremi, Tolulope, et al.
Published: (2023)
WAXAL: A Large-Scale Multilingual African Language Speech Corpus
by: Diack, Abdoulaye, et al.
Published: (2026)
by: Diack, Abdoulaye, et al.
Published: (2026)
CEO: Corpus-based Open-Domain Event Ontology Induction
by: Xu, Nan, et al.
Published: (2023)
by: Xu, Nan, et al.
Published: (2023)
Bundesrecht: An Open Library and Corpus for German Statutory Reference Processing
by: Darji, Harshil, et al.
Published: (2026)
by: Darji, Harshil, et al.
Published: (2026)
StyleBench: Evaluating Speech Language Models on Conversational Speaking Style Control
by: Zhao, Haishu, et al.
Published: (2026)
by: Zhao, Haishu, et al.
Published: (2026)
GlotCC: An Open Broad-Coverage CommonCrawl Corpus and Pipeline for Minority Languages
by: Kargaran, Amir Hossein, et al.
Published: (2024)
by: Kargaran, Amir Hossein, et al.
Published: (2024)
SPOT: An Annotated French Corpus and Benchmark for Detecting Critical Interventions in Online Conversations
by: Berriche, Manon, et al.
Published: (2025)
by: Berriche, Manon, et al.
Published: (2025)
English to Central Kurdish Speech Translation: Corpus Creation, Evaluation, and Orthographic Standardization
by: Mohammadamini, Mohammad, et al.
Published: (2026)
by: Mohammadamini, Mohammad, et al.
Published: (2026)
YouTube-SL-25: A Large-Scale, Open-Domain Multilingual Sign Language Parallel Corpus
by: Tanzer, Garrett, et al.
Published: (2024)
by: Tanzer, Garrett, et al.
Published: (2024)
The TUB Sign Language Corpus Collection
by: Avramidis, Eleftherios, et al.
Published: (2025)
by: Avramidis, Eleftherios, et al.
Published: (2025)
Cross-Lingual Conversational Speech Summarization with Large Language Models
by: Nelson, Max, et al.
Published: (2024)
by: Nelson, Max, et al.
Published: (2024)
Testimole-Conversational: A 30-Billion-Word Italian Discussion Board Corpus (1996-2024) for Language Modeling and Sociolinguistic Research
by: Rinaldi, Matteo, et al.
Published: (2026)
by: Rinaldi, Matteo, et al.
Published: (2026)
WorldSpeech: A Multilingual Speech Corpus from Around the World
by: Asonitis, Antonis, et al.
Published: (2026)
by: Asonitis, Antonis, et al.
Published: (2026)
WenetSpeech-Chuan: A Large-Scale Sichuanese Corpus with Rich Annotation for Dialectal Speech Processing
by: Dai, Yuhang, et al.
Published: (2025)
by: Dai, Yuhang, et al.
Published: (2025)
Omnilingual ASR: Open-Source Multilingual Speech Recognition for 1600+ Languages
by: Omnilingual ASR team, et al.
Published: (2025)
by: Omnilingual ASR team, et al.
Published: (2025)
SafeSpeech: A Comprehensive and Interactive Tool for Analysing Sexist and Abusive Language in Conversations
by: Tan, Xingwei, et al.
Published: (2025)
by: Tan, Xingwei, et al.
Published: (2025)
ESPnet-SpeechLM: An Open Speech Language Model Toolkit
by: Tian, Jinchuan, et al.
Published: (2025)
by: Tian, Jinchuan, et al.
Published: (2025)
OpenWHO: A Document-Level Parallel Corpus for Health Translation in Low-Resource Languages
by: Merx, Raphaël, et al.
Published: (2025)
by: Merx, Raphaël, et al.
Published: (2025)
Ramsa: A Large Sociolinguistically Rich Emirati Arabic Speech Corpus for ASR and TTS
by: Al-Sabbagh, Rania
Published: (2026)
by: Al-Sabbagh, Rania
Published: (2026)
Similar Items
-
ThaiOCRBench: A Task-Diverse Benchmark for Vision-Language Understanding in Thai
by: Nonesung, Surapon, et al.
Published: (2025) -
Typhoon ASR Real-time: FastConformer-Transducer for Thai Automatic Speech Recognition
by: Sirichotedumrong, Warit, et al.
Published: (2026) -
Typhoon 2: A Family of Open Text and Multimodal Thai Large Language Models
by: Pipatanakul, Kunat, et al.
Published: (2024) -
Advancing Speech Translation: A Corpus of Mandarin-English Conversational Telephone Speech
by: Wotherspoon, Shannon, et al.
Published: (2024) -
Speak & Improve Corpus 2025: an L2 English Speech Corpus for Language Assessment and Feedback
by: Knill, Kate, et al.
Published: (2024)