Language Models are Surprisingly Fragile to Drug Names in Biomedical Benchmarks
Fuente:
arXiv
Saved in:
| Main Authors: | Gallifant, Jack, Chen, Shan, Moreira, Pedro, Munch, Nikolaj, Gao, Mingye, Pond, Jackson, Celi, Leo Anthony, Aerts, Hugo, Hartvigsen, Thomas, Bitterman, Danielle |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Wait, but Tylenol is Acetaminophen... Investigating and Improving Language Models' Ability to Resist Requests for Misinformation
by: Chen, Shan, et al.
Published: (2024)
by: Chen, Shan, et al.
Published: (2024)
Sparse Autoencoder Features for Classifications and Transferability
by: Gallifant, Jack, et al.
Published: (2025)
by: Gallifant, Jack, et al.
Published: (2025)
MedBrowseComp: Benchmarking Medical Deep Research and Computer Use
by: Chen, Shan, et al.
Published: (2025)
by: Chen, Shan, et al.
Published: (2025)
Cross-Care: Assessing the Healthcare Implications of Pre-training Data on Language Model Bias
by: Chen, Shan, et al.
Published: (2024)
by: Chen, Shan, et al.
Published: (2024)
KScope: A Framework for Characterizing the Knowledge Status of Language Models
by: Xiao, Yuxin, et al.
Published: (2025)
by: Xiao, Yuxin, et al.
Published: (2025)
debiaSAE: Benchmarking and Mitigating Vision-Language Model Bias
by: Sasse, Kuleen, et al.
Published: (2024)
by: Sasse, Kuleen, et al.
Published: (2024)
Analyzing Diversity in Healthcare LLM Research: A Scientometric Perspective
by: Restrepo, David, et al.
Published: (2024)
by: Restrepo, David, et al.
Published: (2024)
Evaluation of ChatGPT Family of Models for Biomedical Reasoning and Classification
by: Chen, Shan, et al.
Published: (2023)
by: Chen, Shan, et al.
Published: (2023)
Seeds of Stereotypes: A Large-Scale Textual Analysis of Race and Gender Associations with Diseases in Online Sources
by: Hansen, Lasse Hyldig, et al.
Published: (2024)
by: Hansen, Lasse Hyldig, et al.
Published: (2024)
Improving Clinical NLP Performance through Language Model-Generated Synthetic Clinical Data
by: Chen, Shan, et al.
Published: (2024)
by: Chen, Shan, et al.
Published: (2024)
Can Language Models Identify Side Effects of Breast Cancer Radiation Treatments?
by: Seah, Natalie, et al.
Published: (2026)
by: Seah, Natalie, et al.
Published: (2026)
Visually Prompted Benchmarks Are Surprisingly Fragile
by: Feng, Haiwen, et al.
Published: (2025)
by: Feng, Haiwen, et al.
Published: (2025)
WorldMedQA-V: a multilingual, multimodal medical examination dataset for multimodal language models evaluation
by: Matos, João, et al.
Published: (2024)
by: Matos, João, et al.
Published: (2024)
The use of large language models to enhance cancer clinical trial educational materials
by: Gao, Mingye, et al.
Published: (2024)
by: Gao, Mingye, et al.
Published: (2024)
Multi-OphthaLingua: A Multilingual Benchmark for Assessing and Debiasing LLM Ophthalmological QA in LMICs
by: Restrepo, David, et al.
Published: (2024)
by: Restrepo, David, et al.
Published: (2024)
Representation Learning of Lab Values via Masked AutoEncoders
by: Restrepo, David, et al.
Published: (2025)
by: Restrepo, David, et al.
Published: (2025)
Proof of Time: A Benchmark for Evaluating Scientific Idea Judgments
by: Ye, Bingyang, et al.
Published: (2026)
by: Ye, Bingyang, et al.
Published: (2026)
M3: Conversational LLMs Simplify Secure Clinical Data Access, Understanding, and Analysis
by: Attrach, Rafi Al, et al.
Published: (2025)
by: Attrach, Rafi Al, et al.
Published: (2025)
Surprisingly Fragile: Assessing and Addressing Prompt Instability in Multimodal Foundation Models
by: Stewart, Ian, et al.
Published: (2024)
by: Stewart, Ian, et al.
Published: (2024)
Academic Vibe Coding: Opportunities for Accelerating Research in an Era of Resource Constraint
by: Crowson, Matthew G, et al.
Published: (2025)
by: Crowson, Matthew G, et al.
Published: (2025)
Generalization in medical AI: a perspective on developing scalable models
by: Zvuloni, Eran, et al.
Published: (2023)
by: Zvuloni, Eran, et al.
Published: (2023)
When Models Reason in Your Language: Controlling Thinking Language Comes at the Cost of Accuracy
by: Qi, Jirui, et al.
Published: (2025)
by: Qi, Jirui, et al.
Published: (2025)
EHRmonize: A Framework for Medical Concept Abstraction from Electronic Health Records using Large Language Models
by: Matos, João, et al.
Published: (2024)
by: Matos, João, et al.
Published: (2024)
GLiNER-BioMed: A Suite of Efficient Models for Open Biomedical Named Entity Recognition
by: Yazdani, Anthony, et al.
Published: (2025)
by: Yazdani, Anthony, et al.
Published: (2025)
Introductory dynamical oceanography / by Stephen Pond and George L. Pickard
by: Pond, Stephen
Published: (1993)
by: Pond, Stephen
Published: (1993)
Germany's real role in the Ukraine crisis : caught between east and west / Elizabeth Pond, Hans Kundnani
by: Pond, Elizabeth
Published: (2014)
by: Pond, Elizabeth
Published: (2014)
Development of a Professional School Library Association: American Association of School Librarians
by: Pond, Patricia
Published: (1976)
by: Pond, Patricia
Published: (1976)
Large Language Models to Identify Social Determinants of Health in Electronic Health Records
by: Guevara, Marco, et al.
Published: (2023)
by: Guevara, Marco, et al.
Published: (2023)
The Evolution of Lying in a Spatially-Explicit Prisoner's Dilemma Model
by: Hartvigsen, Gregg
Published: (2026)
by: Hartvigsen, Gregg
Published: (2026)
Finding triangle‐free 2‐factors in general graphs
by: David Hartvigsen
Published: (2024)
by: David Hartvigsen
Published: (2024)
The Paradox of Progress: Why Workforce Fragility Remains Nursing's Defining Challenge
by: Debra Jackson
Published: (2026)
by: Debra Jackson
Published: (2026)
Chapter Nanopatterned Surfaces for Biomedical Applications
by: McMurray, Rebecca J., et al.
Published: (2021)
by: McMurray, Rebecca J., et al.
Published: (2021)
Cancer Vaccine Adjuvant Name Recognition from Biomedical Literature using Large Language Models
by: Rehana, Hasin, et al.
Published: (2025)
by: Rehana, Hasin, et al.
Published: (2025)
When Raw Data Prevails: Are Large Language Model Embeddings Effective in Numerical Data Representation for Medical Machine Learning Applications?
by: Gao, Yanjun, et al.
Published: (2024)
by: Gao, Yanjun, et al.
Published: (2024)
Describing additional fluxes to deep sediment traps and water-column decay in a coastal environment
by: Timothy, D., Pond, S
Published: (1991)
by: Timothy, D., Pond, S
Published: (1991)
The Influence of the Spring-neap Tidal Cycle on Currents and Density in Burrard Inlet, British Columbia, Canada
by: Isachsen, P., Pond, S
Published: (1995)
by: Isachsen, P., Pond, S
Published: (1995)
O COMPORTAMENTO SOCIALMENTE RESPONSÁVEL DAS EMPRESAS INFLUENCIA A DECISÃO DE COMPRA DO CONSUMIDOR?
by: Renata Céli Moreira da Silva
Published: (2009)
by: Renata Céli Moreira da Silva
Published: (2009)
GESTÃO INTERNACIONAL: A PRODUÇÃO CIENTÍFICA BRASILEIRA ENTRE 1997 E 2006
by: Renata Céli Moreira da Silva
Published: (2008)
by: Renata Céli Moreira da Silva
Published: (2008)
INTERNACIONALIZAÇÃO DE PEQUENAS EMPRESAS: UM ESTUDO DE CASO COM UMA EMPRESA BRASILEIRA DE TECNOLOGIA
by: Renata Céli Moreira da Silva
Published: (2010)
by: Renata Céli Moreira da Silva
Published: (2010)
OS EFEITOS DO CAUSE-RELATED MARKETING NO COMPORTAMENTO DO CONSUMIDOR
by: Renata Céli Moreira da Silva
Published: (2014)
by: Renata Céli Moreira da Silva
Published: (2014)
Similar Items
-
Wait, but Tylenol is Acetaminophen... Investigating and Improving Language Models' Ability to Resist Requests for Misinformation
by: Chen, Shan, et al.
Published: (2024) -
Sparse Autoencoder Features for Classifications and Transferability
by: Gallifant, Jack, et al.
Published: (2025) -
MedBrowseComp: Benchmarking Medical Deep Research and Computer Use
by: Chen, Shan, et al.
Published: (2025) -
Cross-Care: Assessing the Healthcare Implications of Pre-training Data on Language Model Bias
by: Chen, Shan, et al.
Published: (2024) -
KScope: A Framework for Characterizing the Knowledge Status of Language Models
by: Xiao, Yuxin, et al.
Published: (2025)