Benchmarking and Mitigating Sycophancy in Medical Vision Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xu, Juangui, Guo, Zikun, Lv, Jingwei, Lin, Hongbin, Yang, Shu, Wen, Jun, Wang, Di, Hu, Lijie
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916059909455872
author Xu, Juangui
Guo, Zikun
Lv, Jingwei
Lin, Hongbin
Yang, Shu
Wen, Jun
Wang, Di
Hu, Lijie
author_facet Xu, Juangui
Guo, Zikun
Lv, Jingwei
Lin, Hongbin
Yang, Shu
Wen, Jun
Wang, Di
Hu, Lijie
contents Visual language models (VLMs) have the potential to transform medical workflows. However, the deployment is limited by sycophancy. Despite this serious threat to patient safety, a systematic benchmark remains lacking. This paper addresses this gap by introducing a Medical benchmark that applies multiple templates to VLMs in a hierarchical medical visual question answering task. We find that current VLMs are highly susceptible to visual cues, with failure rates showing a correlation to model size or overall accuracy. we discover that perceived authority and user mimicry are powerful triggers, suggesting a bias mechanism independent of visual data. To overcome this, we propose a Visual Information Purification for Evidence based Responses (VIPER) strategy that proactively filters out non-evidence-based social cues, thereby reinforcing evidence based reasoning. VIPER reduces sycophancy while maintaining interpretability and consistently outperforms baseline methods, laying the necessary foundation for the robust and secure integration of VLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2509_21979
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Benchmarking and Mitigating Sycophancy in Medical Vision Language Models
Xu, Juangui
Guo, Zikun
Lv, Jingwei
Lin, Hongbin
Yang, Shu
Wen, Jun
Wang, Di
Hu, Lijie
Computer Vision and Pattern Recognition
Artificial Intelligence
Visual language models (VLMs) have the potential to transform medical workflows. However, the deployment is limited by sycophancy. Despite this serious threat to patient safety, a systematic benchmark remains lacking. This paper addresses this gap by introducing a Medical benchmark that applies multiple templates to VLMs in a hierarchical medical visual question answering task. We find that current VLMs are highly susceptible to visual cues, with failure rates showing a correlation to model size or overall accuracy. we discover that perceived authority and user mimicry are powerful triggers, suggesting a bias mechanism independent of visual data. To overcome this, we propose a Visual Information Purification for Evidence based Responses (VIPER) strategy that proactively filters out non-evidence-based social cues, thereby reinforcing evidence based reasoning. VIPER reduces sycophancy while maintaining interpretability and consistently outperforms baseline methods, laying the necessary foundation for the robust and secure integration of VLMs.
title Benchmarking and Mitigating Sycophancy in Medical Vision Language Models
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2509.21979