HealthContradict: Evaluating Biomedical Knowledge Conflicts in Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Boya, Bornet, Alban, Yang, Rui, Liu, Nan, Teodoro, Douglas
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911297773240320
author Zhang, Boya
Bornet, Alban
Yang, Rui
Liu, Nan
Teodoro, Douglas
author_facet Zhang, Boya
Bornet, Alban
Yang, Rui
Liu, Nan
Teodoro, Douglas
contents How do language models use contextual information to answer health questions? How are their responses impacted by conflicting contexts? We assess the ability of language models to reason over long, conflicting biomedical contexts using HealthContradict, an expert-verified dataset comprising 920 unique instances, each consisting of a health-related question, a factual answer supported by scientific evidence, and two documents presenting contradictory stances. We consider several prompt settings, including correct, incorrect or contradictory context, and measure their impact on model outputs. Compared to existing medical question-answering evaluation benchmarks, HealthContradict provides greater distinctions of language models' contextual reasoning capabilities. Our experiments show that the strength of fine-tuned biomedical language models lies not only in their parametric knowledge from pretraining, but also in their ability to exploit correct context while resisting incorrect context.
format Preprint
id arxiv_https___arxiv_org_abs_2512_02299
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle HealthContradict: Evaluating Biomedical Knowledge Conflicts in Language Models
Zhang, Boya
Bornet, Alban
Yang, Rui
Liu, Nan
Teodoro, Douglas
Computation and Language
Artificial Intelligence
How do language models use contextual information to answer health questions? How are their responses impacted by conflicting contexts? We assess the ability of language models to reason over long, conflicting biomedical contexts using HealthContradict, an expert-verified dataset comprising 920 unique instances, each consisting of a health-related question, a factual answer supported by scientific evidence, and two documents presenting contradictory stances. We consider several prompt settings, including correct, incorrect or contradictory context, and measure their impact on model outputs. Compared to existing medical question-answering evaluation benchmarks, HealthContradict provides greater distinctions of language models' contextual reasoning capabilities. Our experiments show that the strength of fine-tuned biomedical language models lies not only in their parametric knowledge from pretraining, but also in their ability to exploit correct context while resisting incorrect context.
title HealthContradict: Evaluating Biomedical Knowledge Conflicts in Language Models
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2512.02299