Korean Bio-Medical Corpus (KBMC) for Medical Named Entity Recognition

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Byun, Sungjoo, Hong, Jiseung, Park, Sumin, Jang, Dongjun, Seo, Jean, Kim, Minseok, Oh, Chaeyoung, Shin, Hyopil
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916174612135936
author Byun, Sungjoo
Hong, Jiseung
Park, Sumin
Jang, Dongjun
Seo, Jean
Kim, Minseok
Oh, Chaeyoung
Shin, Hyopil
author_facet Byun, Sungjoo
Hong, Jiseung
Park, Sumin
Jang, Dongjun
Seo, Jean
Kim, Minseok
Oh, Chaeyoung
Shin, Hyopil
contents Named Entity Recognition (NER) plays a pivotal role in medical Natural Language Processing (NLP). Yet, there has not been an open-source medical NER dataset specifically for the Korean language. To address this, we utilized ChatGPT to assist in constructing the KBMC (Korean Bio-Medical Corpus), which we are now presenting to the public. With the KBMC dataset, we noticed an impressive 20% increase in medical NER performance compared to models trained on general Korean NER datasets. This research underscores the significant benefits and importance of using specialized tools and datasets, like ChatGPT, to enhance language processing in specialized fields such as healthcare.
format Preprint
id arxiv_https___arxiv_org_abs_2403_16158
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Korean Bio-Medical Corpus (KBMC) for Medical Named Entity Recognition
Byun, Sungjoo
Hong, Jiseung
Park, Sumin
Jang, Dongjun
Seo, Jean
Kim, Minseok
Oh, Chaeyoung
Shin, Hyopil
Computation and Language
Named Entity Recognition (NER) plays a pivotal role in medical Natural Language Processing (NLP). Yet, there has not been an open-source medical NER dataset specifically for the Korean language. To address this, we utilized ChatGPT to assist in constructing the KBMC (Korean Bio-Medical Corpus), which we are now presenting to the public. With the KBMC dataset, we noticed an impressive 20% increase in medical NER performance compared to models trained on general Korean NER datasets. This research underscores the significant benefits and importance of using specialized tools and datasets, like ChatGPT, to enhance language processing in specialized fields such as healthcare.
title Korean Bio-Medical Corpus (KBMC) for Medical Named Entity Recognition
topic Computation and Language
url https://arxiv.org/abs/2403.16158