LegalMidm: Use-Case-Driven Legal Domain Specialization for Korean Large Language Model

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Jang, Youngjoon, Park, Chanhee, Moon, Hyeonseok, Ham, Young-kyoung, Moon, Jiwon, Kim, Jinhyeon, Jung, JuKyung, Lim, Heuiseok
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914513526194176
author Jang, Youngjoon
Park, Chanhee
Moon, Hyeonseok
Ham, Young-kyoung
Moon, Jiwon
Kim, Jinhyeon
Jung, JuKyung
Lim, Heuiseok
author_facet Jang, Youngjoon
Park, Chanhee
Moon, Hyeonseok
Ham, Young-kyoung
Moon, Jiwon
Kim, Jinhyeon
Jung, JuKyung
Lim, Heuiseok
contents In recent years, the rapid proliferation of open-source large language models (LLMs) has spurred efforts to turn general-purpose models into domain specialists. However, many domain-specialized LLMs are developed using datasets and training protocols that are not aligned with the nuanced requirements of real-world applications. In the legal domain, where precision and reliability are essential, this lack of consideration limits practical utility. In this study, we propose a systematic training framework grounded in the practical needs of the legal domain, with a focus on Korean law. We introduce LegalMidm, a Korean legal-domain LLM, and present a methodology for constructing high-quality, use-case-driven legal datasets and optimized training pipelines. Our approach emphasizes collaboration with legal professionals and rigorous data curation to ensure relevance and factual accuracy, and demonstrates effectiveness in key legal tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2604_25297
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle LegalMidm: Use-Case-Driven Legal Domain Specialization for Korean Large Language Model
Jang, Youngjoon
Park, Chanhee
Moon, Hyeonseok
Ham, Young-kyoung
Moon, Jiwon
Kim, Jinhyeon
Jung, JuKyung
Lim, Heuiseok
Computation and Language
Artificial Intelligence
In recent years, the rapid proliferation of open-source large language models (LLMs) has spurred efforts to turn general-purpose models into domain specialists. However, many domain-specialized LLMs are developed using datasets and training protocols that are not aligned with the nuanced requirements of real-world applications. In the legal domain, where precision and reliability are essential, this lack of consideration limits practical utility. In this study, we propose a systematic training framework grounded in the practical needs of the legal domain, with a focus on Korean law. We introduce LegalMidm, a Korean legal-domain LLM, and present a methodology for constructing high-quality, use-case-driven legal datasets and optimized training pipelines. Our approach emphasizes collaboration with legal professionals and rigorous data curation to ensure relevance and factual accuracy, and demonstrates effectiveness in key legal tasks.
title LegalMidm: Use-Case-Driven Legal Domain Specialization for Korean Large Language Model
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2604.25297