Have the VLMs Lost Confidence? A Study of Sycophancy in VLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Shuo, Ji, Tao, Fan, Xiaoran, Lu, Linsheng, Yang, Leyi, Yang, Yuming, Xi, Zhiheng, Zheng, Rui, Wang, Yuran, Zhao, Xiaohui, Gui, Tao, Zhang, Qi, Huang, Xuanjing
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916439797006336
author Li, Shuo
Ji, Tao
Fan, Xiaoran
Lu, Linsheng
Yang, Leyi
Yang, Yuming
Xi, Zhiheng
Zheng, Rui
Wang, Yuran
Zhao, Xiaohui
Gui, Tao
Zhang, Qi
Huang, Xuanjing
author_facet Li, Shuo
Ji, Tao
Fan, Xiaoran
Lu, Linsheng
Yang, Leyi
Yang, Yuming
Xi, Zhiheng
Zheng, Rui
Wang, Yuran
Zhao, Xiaohui
Gui, Tao
Zhang, Qi
Huang, Xuanjing
contents In the study of LLMs, sycophancy represents a prevalent hallucination that poses significant challenges to these models. Specifically, LLMs often fail to adhere to original correct responses, instead blindly agreeing with users' opinions, even when those opinions are incorrect or malicious. However, research on sycophancy in visual language models (VLMs) has been scarce. In this work, we extend the exploration of sycophancy from LLMs to VLMs, introducing the MM-SY benchmark to evaluate this phenomenon. We present evaluation results from multiple representative models, addressing the gap in sycophancy research for VLMs. To mitigate sycophancy, we propose a synthetic dataset for training and employ methods based on prompts, supervised fine-tuning, and DPO. Our experiments demonstrate that these methods effectively alleviate sycophancy in VLMs. Additionally, we probe VLMs to assess the semantic impact of sycophancy and analyze the attention distribution of visual tokens. Our findings indicate that the ability to prevent sycophancy is predominantly observed in higher layers of the model. The lack of attention to image knowledge in these higher layers may contribute to sycophancy, and enhancing image attention at high layers proves beneficial in mitigating this issue.
format Preprint
id arxiv_https___arxiv_org_abs_2410_11302
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Have the VLMs Lost Confidence? A Study of Sycophancy in VLMs
Li, Shuo
Ji, Tao
Fan, Xiaoran
Lu, Linsheng
Yang, Leyi
Yang, Yuming
Xi, Zhiheng
Zheng, Rui
Wang, Yuran
Zhao, Xiaohui
Gui, Tao
Zhang, Qi
Huang, Xuanjing
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
In the study of LLMs, sycophancy represents a prevalent hallucination that poses significant challenges to these models. Specifically, LLMs often fail to adhere to original correct responses, instead blindly agreeing with users' opinions, even when those opinions are incorrect or malicious. However, research on sycophancy in visual language models (VLMs) has been scarce. In this work, we extend the exploration of sycophancy from LLMs to VLMs, introducing the MM-SY benchmark to evaluate this phenomenon. We present evaluation results from multiple representative models, addressing the gap in sycophancy research for VLMs. To mitigate sycophancy, we propose a synthetic dataset for training and employ methods based on prompts, supervised fine-tuning, and DPO. Our experiments demonstrate that these methods effectively alleviate sycophancy in VLMs. Additionally, we probe VLMs to assess the semantic impact of sycophancy and analyze the attention distribution of visual tokens. Our findings indicate that the ability to prevent sycophancy is predominantly observed in higher layers of the model. The lack of attention to image knowledge in these higher layers may contribute to sycophancy, and enhancing image attention at high layers proves beneficial in mitigating this issue.
title Have the VLMs Lost Confidence? A Study of Sycophancy in VLMs
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2410.11302