Building Korean linguistic resource for NLU data generation of banking app CS dialog system

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yoon, Jeongwoo, Park, On-yu, Hwang, Changhoe, Yoo, Gwanghoon, Laporte, Eric, Nam, Jeesun
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914552455626752
author Yoon, Jeongwoo
Park, On-yu
Hwang, Changhoe
Yoo, Gwanghoon
Laporte, Eric
Nam, Jeesun
author_facet Yoon, Jeongwoo
Park, On-yu
Hwang, Changhoe
Yoo, Gwanghoon
Laporte, Eric
Nam, Jeesun
contents Natural language understanding (NLU) is integral to task-oriented dialog systems, but demands a considerable amount of annotated training data to increase the coverage of diverse utterances. In this study, we report the construction of a linguistic resource named FIAD (Financial Annotated Dataset) and its use to generate a Korean annotated training data for NLU in the banking customer service (CS) domain. By an empirical examination of a corpus of banking app reviews, we identified three linguistic patterns occurring in Korean request utterances: TOPIC (ENTITY, FEATURE), EVENT, and DISCOURSE MARKER. We represented them in LGGs (Local Grammar Graphs) to generate annotated data covering diverse intents and entities. To assess the practicality of the resource, we evaluate the performances of DIET-only (Intent: 0.91 /Topic [entity+feature]: 0.83), DIET+ HANBERT (I:0.94/T:0.85), DIET+ KoBERT (I:0.94/T:0.86), and DIET+ KorBERT (I:0.95/T:0.84) models trained on FIAD-generated data to extract various types of semantic items.
format Preprint
id arxiv_https___arxiv_org_abs_2605_10241
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Building Korean linguistic resource for NLU data generation of banking app CS dialog system
Yoon, Jeongwoo
Park, On-yu
Hwang, Changhoe
Yoo, Gwanghoon
Laporte, Eric
Nam, Jeesun
Computation and Language
Machine Learning
I.7.0
Natural language understanding (NLU) is integral to task-oriented dialog systems, but demands a considerable amount of annotated training data to increase the coverage of diverse utterances. In this study, we report the construction of a linguistic resource named FIAD (Financial Annotated Dataset) and its use to generate a Korean annotated training data for NLU in the banking customer service (CS) domain. By an empirical examination of a corpus of banking app reviews, we identified three linguistic patterns occurring in Korean request utterances: TOPIC (ENTITY, FEATURE), EVENT, and DISCOURSE MARKER. We represented them in LGGs (Local Grammar Graphs) to generate annotated data covering diverse intents and entities. To assess the practicality of the resource, we evaluate the performances of DIET-only (Intent: 0.91 /Topic [entity+feature]: 0.83), DIET+ HANBERT (I:0.94/T:0.85), DIET+ KoBERT (I:0.94/T:0.86), and DIET+ KorBERT (I:0.95/T:0.84) models trained on FIAD-generated data to extract various types of semantic items.
title Building Korean linguistic resource for NLU data generation of banking app CS dialog system
topic Computation and Language
Machine Learning
I.7.0
url https://arxiv.org/abs/2605.10241