Open Korean Historical Corpus: A Millennia-Scale Diachronic Collection of Public Domain Texts

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Song, Seyoung, Kim, Nawon, Chae, Songeun, Park, Kiwoong, Jin, Jiho, Yoo, Haneul, Cho, Kyunghyun, Oh, Alice
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866918476674760704
author Song, Seyoung
Kim, Nawon
Chae, Songeun
Park, Kiwoong
Jin, Jiho
Yoo, Haneul
Cho, Kyunghyun
Oh, Alice
author_facet Song, Seyoung
Kim, Nawon
Chae, Songeun
Park, Kiwoong
Jin, Jiho
Yoo, Haneul
Cho, Kyunghyun
Oh, Alice
contents The history of the Korean language is characterized by a discrepancy between its spoken and written forms and a pivotal shift from Chinese characters to the Hangul alphabet. However, this linguistic evolution has remained largely unexplored in NLP due to a lack of accessible historical corpora. To address this gap, we introduce the Open Korean Historical Corpus, a large-scale, openly licensed dataset spanning 1,300 years and 6 languages, as well as under-represented writing systems like Korean-style Sinitic (Idu) and Hanja-Hangul mixed script. This corpus contains 17.7 million documents and 5.1 billion tokens from 19 sources, ranging from the 7th century to 2025. We leverage this resource to quantitatively analyze major linguistic shifts: (1) Idu usage peaked in the 1860s before declining sharply; (2) the transition from Hanja to Hangul was a rapid transformation starting around 1890; and (3) North Korea's lexical divergence causes modern tokenizers to produce up to 51 times higher out-of-vocabulary rates. This work provides a foundational resource for quantitative diachronic analysis by capturing the history of the Korean language. Moreover, it can serve as a pre-training corpus for large language models, potentially improving their understanding of Sino-Korean vocabulary in modern Hangul as well as archaic writing systems.
format Preprint
id arxiv_https___arxiv_org_abs_2510_24541
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Open Korean Historical Corpus: A Millennia-Scale Diachronic Collection of Public Domain Texts
Song, Seyoung
Kim, Nawon
Chae, Songeun
Park, Kiwoong
Jin, Jiho
Yoo, Haneul
Cho, Kyunghyun
Oh, Alice
Computation and Language
The history of the Korean language is characterized by a discrepancy between its spoken and written forms and a pivotal shift from Chinese characters to the Hangul alphabet. However, this linguistic evolution has remained largely unexplored in NLP due to a lack of accessible historical corpora. To address this gap, we introduce the Open Korean Historical Corpus, a large-scale, openly licensed dataset spanning 1,300 years and 6 languages, as well as under-represented writing systems like Korean-style Sinitic (Idu) and Hanja-Hangul mixed script. This corpus contains 17.7 million documents and 5.1 billion tokens from 19 sources, ranging from the 7th century to 2025. We leverage this resource to quantitatively analyze major linguistic shifts: (1) Idu usage peaked in the 1860s before declining sharply; (2) the transition from Hanja to Hangul was a rapid transformation starting around 1890; and (3) North Korea's lexical divergence causes modern tokenizers to produce up to 51 times higher out-of-vocabulary rates. This work provides a foundational resource for quantitative diachronic analysis by capturing the history of the Korean language. Moreover, it can serve as a pre-training corpus for large language models, potentially improving their understanding of Sino-Korean vocabulary in modern Hangul as well as archaic writing systems.
title Open Korean Historical Corpus: A Millennia-Scale Diachronic Collection of Public Domain Texts
topic Computation and Language
url https://arxiv.org/abs/2510.24541