Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Kim, Jong Myoung, Young-Jun_Lee, Choi, Ho-Jin, Jung, Sangkeun
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:https://arxiv.org/abs/2503.18250
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916662725312512
author Kim, Jong Myoung
Young-Jun_Lee
Choi, Ho-Jin
Jung, Sangkeun
author_facet Kim, Jong Myoung
Young-Jun_Lee
Choi, Ho-Jin
Jung, Sangkeun
contents Transfer learning leverages the abundance of English data to address the scarcity of resources in modeling non-English languages, such as Korean. In this study, we explore the potential of Phrase Aligned Data (PAD) from standardized Statistical Machine Translation (SMT) to enhance the efficiency of transfer learning. Through extensive experiments, we demonstrate that PAD synergizes effectively with the syntactic characteristics of the Korean language, mitigating the weaknesses of SMT and significantly improving model performance. Moreover, we reveal that PAD complements traditional data construction methods and enhances their effectiveness when combined. This innovative approach not only boosts model performance but also suggests a cost-efficient solution for resource-scarce languages.
format Preprint
id arxiv_https___arxiv_org_abs_2503_18250
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle PAD: Towards Efficient Data Generation for Transfer Learning Using Phrase Alignment
Kim, Jong Myoung
Young-Jun_Lee
Choi, Ho-Jin
Jung, Sangkeun
Computation and Language
Transfer learning leverages the abundance of English data to address the scarcity of resources in modeling non-English languages, such as Korean. In this study, we explore the potential of Phrase Aligned Data (PAD) from standardized Statistical Machine Translation (SMT) to enhance the efficiency of transfer learning. Through extensive experiments, we demonstrate that PAD synergizes effectively with the syntactic characteristics of the Korean language, mitigating the weaknesses of SMT and significantly improving model performance. Moreover, we reveal that PAD complements traditional data construction methods and enhances their effectiveness when combined. This innovative approach not only boosts model performance but also suggests a cost-efficient solution for resource-scarce languages.
title PAD: Towards Efficient Data Generation for Transfer Learning Using Phrase Alignment
topic Computation and Language
url https://arxiv.org/abs/2503.18250