LegalMidm: Use-Case-Driven Legal Domain Specialization for Korean Large Language Model
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866914513526194176 |
|---|---|
| author | Jang, Youngjoon Park, Chanhee Moon, Hyeonseok Ham, Young-kyoung Moon, Jiwon Kim, Jinhyeon Jung, JuKyung Lim, Heuiseok |
| author_facet | Jang, Youngjoon Park, Chanhee Moon, Hyeonseok Ham, Young-kyoung Moon, Jiwon Kim, Jinhyeon Jung, JuKyung Lim, Heuiseok |
| contents | In recent years, the rapid proliferation of open-source large language models (LLMs) has spurred efforts to turn general-purpose models into domain specialists. However, many domain-specialized LLMs are developed using datasets and training protocols that are not aligned with the nuanced requirements of real-world applications. In the legal domain, where precision and reliability are essential, this lack of consideration limits practical utility. In this study, we propose a systematic training framework grounded in the practical needs of the legal domain, with a focus on Korean law. We introduce LegalMidm, a Korean legal-domain LLM, and present a methodology for constructing high-quality, use-case-driven legal datasets and optimized training pipelines. Our approach emphasizes collaboration with legal professionals and rigorous data curation to ensure relevance and factual accuracy, and demonstrates effectiveness in key legal tasks. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2604_25297 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | LegalMidm: Use-Case-Driven Legal Domain Specialization for Korean Large Language Model Jang, Youngjoon Park, Chanhee Moon, Hyeonseok Ham, Young-kyoung Moon, Jiwon Kim, Jinhyeon Jung, JuKyung Lim, Heuiseok Computation and Language Artificial Intelligence In recent years, the rapid proliferation of open-source large language models (LLMs) has spurred efforts to turn general-purpose models into domain specialists. However, many domain-specialized LLMs are developed using datasets and training protocols that are not aligned with the nuanced requirements of real-world applications. In the legal domain, where precision and reliability are essential, this lack of consideration limits practical utility. In this study, we propose a systematic training framework grounded in the practical needs of the legal domain, with a focus on Korean law. We introduce LegalMidm, a Korean legal-domain LLM, and present a methodology for constructing high-quality, use-case-driven legal datasets and optimized training pipelines. Our approach emphasizes collaboration with legal professionals and rigorous data curation to ensure relevance and factual accuracy, and demonstrates effectiveness in key legal tasks. |
| title | LegalMidm: Use-Case-Driven Legal Domain Specialization for Korean Large Language Model |
| topic | Computation and Language Artificial Intelligence |
| url | https://arxiv.org/abs/2604.25297 |