Language Models are Surprisingly Fragile to Drug Names in Biomedical Benchmarks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gallifant, Jack, Chen, Shan, Moreira, Pedro, Munch, Nikolaj, Gao, Mingye, Pond, Jackson, Celi, Leo Anthony, Aerts, Hugo, Hartvigsen, Thomas, Bitterman, Danielle
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913397998616576
author Gallifant, Jack
Chen, Shan
Moreira, Pedro
Munch, Nikolaj
Gao, Mingye
Pond, Jackson
Celi, Leo Anthony
Aerts, Hugo
Hartvigsen, Thomas
Bitterman, Danielle
author_facet Gallifant, Jack
Chen, Shan
Moreira, Pedro
Munch, Nikolaj
Gao, Mingye
Pond, Jackson
Celi, Leo Anthony
Aerts, Hugo
Hartvigsen, Thomas
Bitterman, Danielle
contents Medical knowledge is context-dependent and requires consistent reasoning across various natural language expressions of semantically equivalent phrases. This is particularly crucial for drug names, where patients often use brand names like Advil or Tylenol instead of their generic equivalents. To study this, we create a new robustness dataset, RABBITS, to evaluate performance differences on medical benchmarks after swapping brand and generic drug names using physician expert annotations. We assess both open-source and API-based LLMs on MedQA and MedMCQA, revealing a consistent performance drop ranging from 1-10\%. Furthermore, we identify a potential source of this fragility as the contamination of test data in widely used pre-training datasets. All code is accessible at https://github.com/BittermanLab/RABBITS, and a HuggingFace leaderboard is available at https://huggingface.co/spaces/AIM-Harvard/rabbits-leaderboard.
format Preprint
id arxiv_https___arxiv_org_abs_2406_12066
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Language Models are Surprisingly Fragile to Drug Names in Biomedical Benchmarks
Gallifant, Jack
Chen, Shan
Moreira, Pedro
Munch, Nikolaj
Gao, Mingye
Pond, Jackson
Celi, Leo Anthony
Aerts, Hugo
Hartvigsen, Thomas
Bitterman, Danielle
Computation and Language
Medical knowledge is context-dependent and requires consistent reasoning across various natural language expressions of semantically equivalent phrases. This is particularly crucial for drug names, where patients often use brand names like Advil or Tylenol instead of their generic equivalents. To study this, we create a new robustness dataset, RABBITS, to evaluate performance differences on medical benchmarks after swapping brand and generic drug names using physician expert annotations. We assess both open-source and API-based LLMs on MedQA and MedMCQA, revealing a consistent performance drop ranging from 1-10\%. Furthermore, we identify a potential source of this fragility as the contamination of test data in widely used pre-training datasets. All code is accessible at https://github.com/BittermanLab/RABBITS, and a HuggingFace leaderboard is available at https://huggingface.co/spaces/AIM-Harvard/rabbits-leaderboard.
title Language Models are Surprisingly Fragile to Drug Names in Biomedical Benchmarks
topic Computation and Language
url https://arxiv.org/abs/2406.12066