Language Models are Surprisingly Fragile to Drug Names in Biomedical Benchmarks
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866913397998616576 |
|---|---|
| author | Gallifant, Jack Chen, Shan Moreira, Pedro Munch, Nikolaj Gao, Mingye Pond, Jackson Celi, Leo Anthony Aerts, Hugo Hartvigsen, Thomas Bitterman, Danielle |
| author_facet | Gallifant, Jack Chen, Shan Moreira, Pedro Munch, Nikolaj Gao, Mingye Pond, Jackson Celi, Leo Anthony Aerts, Hugo Hartvigsen, Thomas Bitterman, Danielle |
| contents | Medical knowledge is context-dependent and requires consistent reasoning across various natural language expressions of semantically equivalent phrases. This is particularly crucial for drug names, where patients often use brand names like Advil or Tylenol instead of their generic equivalents. To study this, we create a new robustness dataset, RABBITS, to evaluate performance differences on medical benchmarks after swapping brand and generic drug names using physician expert annotations.
We assess both open-source and API-based LLMs on MedQA and MedMCQA, revealing a consistent performance drop ranging from 1-10\%. Furthermore, we identify a potential source of this fragility as the contamination of test data in widely used pre-training datasets. All code is accessible at https://github.com/BittermanLab/RABBITS, and a HuggingFace leaderboard is available at https://huggingface.co/spaces/AIM-Harvard/rabbits-leaderboard. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2406_12066 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Language Models are Surprisingly Fragile to Drug Names in Biomedical Benchmarks Gallifant, Jack Chen, Shan Moreira, Pedro Munch, Nikolaj Gao, Mingye Pond, Jackson Celi, Leo Anthony Aerts, Hugo Hartvigsen, Thomas Bitterman, Danielle Computation and Language Medical knowledge is context-dependent and requires consistent reasoning across various natural language expressions of semantically equivalent phrases. This is particularly crucial for drug names, where patients often use brand names like Advil or Tylenol instead of their generic equivalents. To study this, we create a new robustness dataset, RABBITS, to evaluate performance differences on medical benchmarks after swapping brand and generic drug names using physician expert annotations. We assess both open-source and API-based LLMs on MedQA and MedMCQA, revealing a consistent performance drop ranging from 1-10\%. Furthermore, we identify a potential source of this fragility as the contamination of test data in widely used pre-training datasets. All code is accessible at https://github.com/BittermanLab/RABBITS, and a HuggingFace leaderboard is available at https://huggingface.co/spaces/AIM-Harvard/rabbits-leaderboard. |
| title | Language Models are Surprisingly Fragile to Drug Names in Biomedical Benchmarks |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2406.12066 |