Guardado en:
Detalles Bibliográficos
Autores principales: Wu, Weiyi, Xu, Xinwen, Gao, Chongyang, Diao, Xingjian, Li, Siting, Salas, Lucas A., Gui, Jiang
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:https://arxiv.org/abs/2505.07968
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912573739237376
author Wu, Weiyi
Xu, Xinwen
Gao, Chongyang
Diao, Xingjian
Li, Siting
Salas, Lucas A.
Gui, Jiang
author_facet Wu, Weiyi
Xu, Xinwen
Gao, Chongyang
Diao, Xingjian
Li, Siting
Salas, Lucas A.
Gui, Jiang
contents Large Language Models (LLMs) have great potential in the field of health care, yet they face great challenges in adapting to rapidly evolving medical knowledge. This can lead to outdated or contradictory treatment suggestions. This study investigated how LLMs respond to evolving clinical guidelines, focusing on concept drift and internal inconsistencies. We developed the DriftMedQA benchmark to simulate guideline evolution and assessed the temporal reliability of various LLMs. Our evaluation of seven state-of-the-art models across 4,290 scenarios demonstrated difficulties in rejecting outdated recommendations and frequently endorsing conflicting guidance. Additionally, we explored two mitigation strategies: Retrieval-Augmented Generation and preference fine-tuning via Direct Preference Optimization. While each method improved model performance, their combination led to the most consistent and reliable results. These findings underscore the need to improve LLM robustness to temporal shifts to ensure more dependable applications in clinical practice. The dataset is available at https://huggingface.co/datasets/RDBH/DriftMed.
format Preprint
id arxiv_https___arxiv_org_abs_2505_07968
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Assessing and Mitigating Medical Knowledge Drift and Conflicts in Large Language Models
Wu, Weiyi
Xu, Xinwen
Gao, Chongyang
Diao, Xingjian
Li, Siting
Salas, Lucas A.
Gui, Jiang
Computation and Language
Large Language Models (LLMs) have great potential in the field of health care, yet they face great challenges in adapting to rapidly evolving medical knowledge. This can lead to outdated or contradictory treatment suggestions. This study investigated how LLMs respond to evolving clinical guidelines, focusing on concept drift and internal inconsistencies. We developed the DriftMedQA benchmark to simulate guideline evolution and assessed the temporal reliability of various LLMs. Our evaluation of seven state-of-the-art models across 4,290 scenarios demonstrated difficulties in rejecting outdated recommendations and frequently endorsing conflicting guidance. Additionally, we explored two mitigation strategies: Retrieval-Augmented Generation and preference fine-tuning via Direct Preference Optimization. While each method improved model performance, their combination led to the most consistent and reliable results. These findings underscore the need to improve LLM robustness to temporal shifts to ensure more dependable applications in clinical practice. The dataset is available at https://huggingface.co/datasets/RDBH/DriftMed.
title Assessing and Mitigating Medical Knowledge Drift and Conflicts in Large Language Models
topic Computation and Language
url https://arxiv.org/abs/2505.07968