When Prompts Override Vision: Prompt-Induced Hallucinations in LVLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Khayatan, Pegah, Parekh, Jayneel, Dapogny, Arnaud, Shukor, Mustafa, Newson, Alasdair, Cord, Matthieu
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908989882630144
author Khayatan, Pegah
Parekh, Jayneel
Dapogny, Arnaud
Shukor, Mustafa
Newson, Alasdair
Cord, Matthieu
author_facet Khayatan, Pegah
Parekh, Jayneel
Dapogny, Arnaud
Shukor, Mustafa
Newson, Alasdair
Cord, Matthieu
contents Despite impressive progress in capabilities of large vision-language models (LVLMs), these systems remain vulnerable to hallucinations, i.e., outputs that are not grounded in the visual input. Prior work has attributed hallucinations in LVLMs to factors such as limitations of the vision backbone or the dominance of the language component, yet the relative importance of these factors remains unclear. To resolve this ambiguity, We propose HalluScope, a benchmark to better understand the extent to which different factors induce hallucinations. Our analysis indicates that hallucinations largely stem from excessive reliance on textual priors and background knowledge, especially information introduced through textual instructions. To mitigate hallucinations induced by textual instruction priors, we propose HalluVL-DPO, a framework for fine-tuning off-the-shelf LVLMs towards more visually grounded responses. HalluVL-DPO leverages preference optimization using a curated training dataset that we construct, guiding the model to prefer grounded responses over hallucinated ones. We demonstrate that our optimized model effectively mitigates the targeted hallucination failure mode, while preserving or improving performance on other hallucination benchmarks and visual capability evaluations. To support reproducibility and further research, we will publicly release our evaluation benchmark, preference training dataset, and code at https://pegah-kh.github.io/projects/prompts-override-vision/ .
format Preprint
id arxiv_https___arxiv_org_abs_2604_21911
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle When Prompts Override Vision: Prompt-Induced Hallucinations in LVLMs
Khayatan, Pegah
Parekh, Jayneel
Dapogny, Arnaud
Shukor, Mustafa
Newson, Alasdair
Cord, Matthieu
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
Despite impressive progress in capabilities of large vision-language models (LVLMs), these systems remain vulnerable to hallucinations, i.e., outputs that are not grounded in the visual input. Prior work has attributed hallucinations in LVLMs to factors such as limitations of the vision backbone or the dominance of the language component, yet the relative importance of these factors remains unclear. To resolve this ambiguity, We propose HalluScope, a benchmark to better understand the extent to which different factors induce hallucinations. Our analysis indicates that hallucinations largely stem from excessive reliance on textual priors and background knowledge, especially information introduced through textual instructions. To mitigate hallucinations induced by textual instruction priors, we propose HalluVL-DPO, a framework for fine-tuning off-the-shelf LVLMs towards more visually grounded responses. HalluVL-DPO leverages preference optimization using a curated training dataset that we construct, guiding the model to prefer grounded responses over hallucinated ones. We demonstrate that our optimized model effectively mitigates the targeted hallucination failure mode, while preserving or improving performance on other hallucination benchmarks and visual capability evaluations. To support reproducibility and further research, we will publicly release our evaluation benchmark, preference training dataset, and code at https://pegah-kh.github.io/projects/prompts-override-vision/ .
title When Prompts Override Vision: Prompt-Induced Hallucinations in LVLMs
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2604.21911