Pointing to a Llama and Call it a Camel: On the Sycophancy of Multimodal Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Pi, Renjie, Miao, Kehao, Peihang, Li, Liu, Runtao, Gao, Jiahui, Zhang, Jipeng, Zhou, Xiaofang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914047434162176
author Pi, Renjie
Miao, Kehao
Peihang, Li
Liu, Runtao
Gao, Jiahui
Zhang, Jipeng
Zhou, Xiaofang
author_facet Pi, Renjie
Miao, Kehao
Peihang, Li
Liu, Runtao
Gao, Jiahui
Zhang, Jipeng
Zhou, Xiaofang
contents Multimodal large language models (MLLMs) have demonstrated extraordinary capabilities in conducting conversations based on image inputs. However, we observe that MLLMs exhibit a pronounced form of visual sycophantic behavior. While similar behavior has also been noted in text-based large language models (LLMs), it becomes significantly more prominent when MLLMs process image inputs. We refer to this phenomenon as the "sycophantic modality gap." To better understand this issue, we further analyze the factors that contribute to the exacerbation of this gap. To mitigate the visual sycophantic behavior, we first experiment with naive supervised fine-tuning to help the MLLM resist misleading instructions from the user. However, we find that this approach also makes the MLLM overly resistant to corrective instructions (i.e., stubborn even if it is wrong). To alleviate this trade-off, we propose Sycophantic Reflective Tuning (SRT), which enables the MLLM to engage in reflective reasoning, allowing it to determine whether a user's instruction is misleading or corrective before drawing a conclusion. After applying SRT, we observe a significant reduction in sycophantic behavior toward misleading instructions, without resulting in excessive stubbornness when receiving corrective instructions.
format Preprint
id arxiv_https___arxiv_org_abs_2509_16149
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Pointing to a Llama and Call it a Camel: On the Sycophancy of Multimodal Large Language Models
Pi, Renjie
Miao, Kehao
Peihang, Li
Liu, Runtao
Gao, Jiahui
Zhang, Jipeng
Zhou, Xiaofang
Computer Vision and Pattern Recognition
Multimodal large language models (MLLMs) have demonstrated extraordinary capabilities in conducting conversations based on image inputs. However, we observe that MLLMs exhibit a pronounced form of visual sycophantic behavior. While similar behavior has also been noted in text-based large language models (LLMs), it becomes significantly more prominent when MLLMs process image inputs. We refer to this phenomenon as the "sycophantic modality gap." To better understand this issue, we further analyze the factors that contribute to the exacerbation of this gap. To mitigate the visual sycophantic behavior, we first experiment with naive supervised fine-tuning to help the MLLM resist misleading instructions from the user. However, we find that this approach also makes the MLLM overly resistant to corrective instructions (i.e., stubborn even if it is wrong). To alleviate this trade-off, we propose Sycophantic Reflective Tuning (SRT), which enables the MLLM to engage in reflective reasoning, allowing it to determine whether a user's instruction is misleading or corrective before drawing a conclusion. After applying SRT, we observe a significant reduction in sycophantic behavior toward misleading instructions, without resulting in excessive stubbornness when receiving corrective instructions.
title Pointing to a Llama and Call it a Camel: On the Sycophancy of Multimodal Large Language Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.16149