Medical Coding with Biomedical Transformer Ensembles and Zero/Few-shot Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ziletti, Angelo, Akbik, Alan, Berns, Christoph, Herold, Thomas, Legler, Marion, Viell, Martina
Format: Preprint
Published: 2022
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913638185435136
author Ziletti, Angelo
Akbik, Alan
Berns, Christoph
Herold, Thomas
Legler, Marion
Viell, Martina
author_facet Ziletti, Angelo
Akbik, Alan
Berns, Christoph
Herold, Thomas
Legler, Marion
Viell, Martina
contents Medical coding (MC) is an essential pre-requisite for reliable data retrieval and reporting. Given a free-text reported term (RT) such as "pain of right thigh to the knee", the task is to identify the matching lowest-level term (LLT) - in this case "unilateral leg pain" - from a very large and continuously growing repository of standardized medical terms. However, automating this task is challenging due to a large number of LLT codes (as of writing over 80,000), limited availability of training data for long tail/emerging classes, and the general high accuracy demands of the medical domain. With this paper, we introduce the MC task, discuss its challenges, and present a novel approach called xTARS that combines traditional BERT-based classification with a recent zero/few-shot learning approach (TARS). We present extensive experiments that show that our combined approach outperforms strong baselines, especially in the few-shot regime. The approach is developed and deployed at Bayer, live since November 2021. As we believe our approach potentially promising beyond MC, and to ensure reproducibility, we release the code to the research community.
format Preprint
id arxiv_https___arxiv_org_abs_2206_02662
institution arXiv
publishDate 2022
record_format arxiv
spellingShingle Medical Coding with Biomedical Transformer Ensembles and Zero/Few-shot Learning
Ziletti, Angelo
Akbik, Alan
Berns, Christoph
Herold, Thomas
Legler, Marion
Viell, Martina
Information Retrieval
Computation and Language
Machine Learning
Medical coding (MC) is an essential pre-requisite for reliable data retrieval and reporting. Given a free-text reported term (RT) such as "pain of right thigh to the knee", the task is to identify the matching lowest-level term (LLT) - in this case "unilateral leg pain" - from a very large and continuously growing repository of standardized medical terms. However, automating this task is challenging due to a large number of LLT codes (as of writing over 80,000), limited availability of training data for long tail/emerging classes, and the general high accuracy demands of the medical domain. With this paper, we introduce the MC task, discuss its challenges, and present a novel approach called xTARS that combines traditional BERT-based classification with a recent zero/few-shot learning approach (TARS). We present extensive experiments that show that our combined approach outperforms strong baselines, especially in the few-shot regime. The approach is developed and deployed at Bayer, live since November 2021. As we believe our approach potentially promising beyond MC, and to ensure reproducibility, we release the code to the research community.
title Medical Coding with Biomedical Transformer Ensembles and Zero/Few-shot Learning
topic Information Retrieval
Computation and Language
Machine Learning
url https://arxiv.org/abs/2206.02662