Saved in:
Bibliographic Details
Main Authors: Yun, Hye Sun, Zhang, Karen Y. C., Kouzy, Ramez, Marshall, Iain J., Li, Junyi Jessy, Wallace, Byron C.
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2502.07963
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911613270884352
author Yun, Hye Sun
Zhang, Karen Y. C.
Kouzy, Ramez
Marshall, Iain J.
Li, Junyi Jessy
Wallace, Byron C.
author_facet Yun, Hye Sun
Zhang, Karen Y. C.
Kouzy, Ramez
Marshall, Iain J.
Li, Junyi Jessy
Wallace, Byron C.
contents Medical research faces well-documented challenges in translating novel treatments into clinical practice. Publishing incentives encourage researchers to present "positive" findings, even when empirical results are equivocal. Consequently, it is well-documented that authors often spin study results, especially in article abstracts. Such spin can influence clinician interpretation of evidence and may affect patient care decisions. In this study, we ask whether the interpretation of trial results offered by Large Language Models (LLMs) is similarly affected by spin. This is important since LLMs are increasingly being used to trawl through and synthesize published medical evidence. We evaluated 22 LLMs and found that they are across the board more susceptible to spin than humans. They might also propagate spin into their outputs: We find evidence, e.g., that LLMs implicitly incorporate spin into plain language summaries that they generate. We also find, however, that LLMs are generally capable of recognizing spin, and can be prompted in a way to mitigate spin's impact on LLM outputs.
format Preprint
id arxiv_https___arxiv_org_abs_2502_07963
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Caught in the Web of Words: Do LLMs Fall for Spin in Medical Literature?
Yun, Hye Sun
Zhang, Karen Y. C.
Kouzy, Ramez
Marshall, Iain J.
Li, Junyi Jessy
Wallace, Byron C.
Computation and Language
Artificial Intelligence
Medical research faces well-documented challenges in translating novel treatments into clinical practice. Publishing incentives encourage researchers to present "positive" findings, even when empirical results are equivocal. Consequently, it is well-documented that authors often spin study results, especially in article abstracts. Such spin can influence clinician interpretation of evidence and may affect patient care decisions. In this study, we ask whether the interpretation of trial results offered by Large Language Models (LLMs) is similarly affected by spin. This is important since LLMs are increasingly being used to trawl through and synthesize published medical evidence. We evaluated 22 LLMs and found that they are across the board more susceptible to spin than humans. They might also propagate spin into their outputs: We find evidence, e.g., that LLMs implicitly incorporate spin into plain language summaries that they generate. We also find, however, that LLMs are generally capable of recognizing spin, and can be prompted in a way to mitigate spin's impact on LLM outputs.
title Caught in the Web of Words: Do LLMs Fall for Spin in Medical Literature?
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2502.07963