DOTA-ME-CS: Daily Oriented Text Audio-Mandarin English-Code Switching Dataset

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Yupei, Wei, Zifan, Yu, Heng, Xue, Jiahao, Zhou, Huichi, Schuller, Björn W.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908649178267648
author Li, Yupei
Wei, Zifan
Yu, Heng
Xue, Jiahao
Zhou, Huichi
Schuller, Björn W.
author_facet Li, Yupei
Wei, Zifan
Yu, Heng
Xue, Jiahao
Zhou, Huichi
Schuller, Björn W.
contents Code-switching, the alternation between two or more languages within communication, poses great challenges for Automatic Speech Recognition (ASR) systems. Existing models and datasets are limited in their ability to effectively handle these challenges. To address this gap and foster progress in code-switching ASR research, we introduce the DOTA-ME-CS: Daily oriented text audio Mandarin-English code-switching dataset, which consists of 18.54 hours of audio data, including 9,300 recordings from 34 participants. To enhance the dataset's diversity, we apply artificial intelligence (AI) techniques such as AI timbre synthesis, speed variation, and noise addition, thereby increasing the complexity and scalability of the task. The dataset is carefully curated to ensure both diversity and quality, providing a robust resource for researchers addressing the intricacies of bilingual speech recognition with detailed data analysis. We further demonstrate the dataset's potential in future research. The DOTA-ME-CS dataset, along with accompanying code, will be made publicly available.
format Preprint
id arxiv_https___arxiv_org_abs_2501_12122
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DOTA-ME-CS: Daily Oriented Text Audio-Mandarin English-Code Switching Dataset
Li, Yupei
Wei, Zifan
Yu, Heng
Xue, Jiahao
Zhou, Huichi
Schuller, Björn W.
Sound
Audio and Speech Processing
Code-switching, the alternation between two or more languages within communication, poses great challenges for Automatic Speech Recognition (ASR) systems. Existing models and datasets are limited in their ability to effectively handle these challenges. To address this gap and foster progress in code-switching ASR research, we introduce the DOTA-ME-CS: Daily oriented text audio Mandarin-English code-switching dataset, which consists of 18.54 hours of audio data, including 9,300 recordings from 34 participants. To enhance the dataset's diversity, we apply artificial intelligence (AI) techniques such as AI timbre synthesis, speed variation, and noise addition, thereby increasing the complexity and scalability of the task. The dataset is carefully curated to ensure both diversity and quality, providing a robust resource for researchers addressing the intricacies of bilingual speech recognition with detailed data analysis. We further demonstrate the dataset's potential in future research. The DOTA-ME-CS dataset, along with accompanying code, will be made publicly available.
title DOTA-ME-CS: Daily Oriented Text Audio-Mandarin English-Code Switching Dataset
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2501.12122