Enregistré dans:
Détails bibliographiques
Auteurs principaux: Cui, Wenqian, Fu, Xiangling, Liu, Shaohui, Gu, Mingjun, Liu, Xien, Wu, Ji, King, Irwin
Format: Preprint
Publié: 2023
Sujets:
Accès en ligne:https://arxiv.org/abs/2306.01931
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866917691898462208
author Cui, Wenqian
Fu, Xiangling
Liu, Shaohui
Gu, Mingjun
Liu, Xien
Wu, Ji
King, Irwin
author_facet Cui, Wenqian
Fu, Xiangling
Liu, Shaohui
Gu, Mingjun
Liu, Xien
Wu, Ji
King, Irwin
contents Disease name normalization is an important task in the medical domain. It classifies disease names written in various formats into standardized names, serving as a fundamental component in smart healthcare systems for various disease-related functions. Nevertheless, the most significant obstacle to existing disease name normalization systems is the severe shortage of training data. Consequently, we present a novel data augmentation approach that includes a series of data augmentation techniques and some supporting modules to help mitigate the problem. Our proposed methods rely on the Structural Invariance property of disease names and the Hierarchy property of the disease classification system. The goal is to equip the models with extensive understanding of the disease names and the hierarchical structure of the disease name classification system. Through extensive experimentation, we illustrate that our proposed approach exhibits significant performance improvements across various baseline models and training objectives, particularly in scenarios with limited training data.
format Preprint
id arxiv_https___arxiv_org_abs_2306_01931
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Simple Data Augmentation Techniques for Chinese Disease Normalization
Cui, Wenqian
Fu, Xiangling
Liu, Shaohui
Gu, Mingjun
Liu, Xien
Wu, Ji
King, Irwin
Computation and Language
Artificial Intelligence
Disease name normalization is an important task in the medical domain. It classifies disease names written in various formats into standardized names, serving as a fundamental component in smart healthcare systems for various disease-related functions. Nevertheless, the most significant obstacle to existing disease name normalization systems is the severe shortage of training data. Consequently, we present a novel data augmentation approach that includes a series of data augmentation techniques and some supporting modules to help mitigate the problem. Our proposed methods rely on the Structural Invariance property of disease names and the Hierarchy property of the disease classification system. The goal is to equip the models with extensive understanding of the disease names and the hierarchical structure of the disease name classification system. Through extensive experimentation, we illustrate that our proposed approach exhibits significant performance improvements across various baseline models and training objectives, particularly in scenarios with limited training data.
title Simple Data Augmentation Techniques for Chinese Disease Normalization
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2306.01931