HiKE: Hierarchical Evaluation Framework for Korean-English Code-Switching Speech Recognition

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Paik, Gio, Kim, Yongbeom, Lee, Soungmin, Ahn, Sangmin, Kim, Chanwoo
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909988239179776
author Paik, Gio
Kim, Yongbeom
Lee, Soungmin
Ahn, Sangmin
Kim, Chanwoo
author_facet Paik, Gio
Kim, Yongbeom
Lee, Soungmin
Ahn, Sangmin
Kim, Chanwoo
contents Despite advances in multilingual automatic speech recognition (ASR), code-switching (CS), the mixing of languages within an utterance common in daily speech, remains a severely underexplored challenge. In this paper, we introduce HiKE: the Hierarchical Korean-English code-switching benchmark, the first globally accessible non-synthetic evaluation framework for Korean-English CS, aiming to provide a means for the precise evaluation of multilingual ASR models and to foster research in the field. The proposed framework not only consists of high-quality, natural CS data across various topics, but also provides meticulous loanword labels and a hierarchical CS-level labeling scheme (word, phrase, and sentence) that together enable a systematic evaluation of a model's ability to handle each distinct level of code-switching. Through evaluations of diverse multilingual ASR models and fine-tuning experiments, this paper demonstrates that although most multilingual ASR models initially exhibit inadequate CS-ASR performance, this capability can be enabled through fine-tuning with synthetic CS data. HiKE is available at https://github.com/ThetaOne-AI/HiKE.
format Preprint
id arxiv_https___arxiv_org_abs_2509_24613
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle HiKE: Hierarchical Evaluation Framework for Korean-English Code-Switching Speech Recognition
Paik, Gio
Kim, Yongbeom
Lee, Soungmin
Ahn, Sangmin
Kim, Chanwoo
Computation and Language
Sound
Audio and Speech Processing
Despite advances in multilingual automatic speech recognition (ASR), code-switching (CS), the mixing of languages within an utterance common in daily speech, remains a severely underexplored challenge. In this paper, we introduce HiKE: the Hierarchical Korean-English code-switching benchmark, the first globally accessible non-synthetic evaluation framework for Korean-English CS, aiming to provide a means for the precise evaluation of multilingual ASR models and to foster research in the field. The proposed framework not only consists of high-quality, natural CS data across various topics, but also provides meticulous loanword labels and a hierarchical CS-level labeling scheme (word, phrase, and sentence) that together enable a systematic evaluation of a model's ability to handle each distinct level of code-switching. Through evaluations of diverse multilingual ASR models and fine-tuning experiments, this paper demonstrates that although most multilingual ASR models initially exhibit inadequate CS-ASR performance, this capability can be enabled through fine-tuning with synthetic CS data. HiKE is available at https://github.com/ThetaOne-AI/HiKE.
title HiKE: Hierarchical Evaluation Framework for Korean-English Code-Switching Speech Recognition
topic Computation and Language
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2509.24613