Developing an Open Conversational Speech Corpus for the Isan Language

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Na-Thalang, Adisai, Wittayasakpan, Chanakan, Phatcharoen, Kritsadha, Buakaw, Supakit
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912747839553536
author Na-Thalang, Adisai
Wittayasakpan, Chanakan
Phatcharoen, Kritsadha
Buakaw, Supakit
author_facet Na-Thalang, Adisai
Wittayasakpan, Chanakan
Phatcharoen, Kritsadha
Buakaw, Supakit
contents This paper introduces the development of the first open conversational speech dataset for the Isan language, the most widely spoken regional dialect in Thailand. Unlike existing speech corpora that are primarily based on read or scripted speech, this dataset consists of natural speech, thereby capturing authentic linguistic phenomena such as colloquials, spontaneous prosody, disfluencies, and frequent code-switching with central Thai. A key challenge in building this resource lies in the lack of a standardized orthography for Isan. Current writing practices vary considerably, due to the different lexical tones between Thai and Isan. This variability complicates the design of transcription guidelines and poses questions regarding consistency, usability, and linguistic authenticity. To address these issues, we establish practical transcription protocols that balance the need for representational accuracy with the requirements of computational processing. By releasing this dataset as an open resource, we aim to contribute to inclusive AI development, support research on underrepresented languages, and provide a basis for addressing the linguistic and technical challenges inherent in modeling conversational speech.
format Preprint
id arxiv_https___arxiv_org_abs_2511_21229
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Developing an Open Conversational Speech Corpus for the Isan Language
Na-Thalang, Adisai
Wittayasakpan, Chanakan
Phatcharoen, Kritsadha
Buakaw, Supakit
Computation and Language
This paper introduces the development of the first open conversational speech dataset for the Isan language, the most widely spoken regional dialect in Thailand. Unlike existing speech corpora that are primarily based on read or scripted speech, this dataset consists of natural speech, thereby capturing authentic linguistic phenomena such as colloquials, spontaneous prosody, disfluencies, and frequent code-switching with central Thai. A key challenge in building this resource lies in the lack of a standardized orthography for Isan. Current writing practices vary considerably, due to the different lexical tones between Thai and Isan. This variability complicates the design of transcription guidelines and poses questions regarding consistency, usability, and linguistic authenticity. To address these issues, we establish practical transcription protocols that balance the need for representational accuracy with the requirements of computational processing. By releasing this dataset as an open resource, we aim to contribute to inclusive AI development, support research on underrepresented languages, and provide a basis for addressing the linguistic and technical challenges inherent in modeling conversational speech.
title Developing an Open Conversational Speech Corpus for the Isan Language
topic Computation and Language
url https://arxiv.org/abs/2511.21229