Black-Box Visual Prompt Engineering for Mitigating Object Hallucination in Large Vision Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Woo, Sangmin, Zhou, Kang, Zhou, Yun, Wang, Shuai, Guan, Sheng, Ding, Haibo, Cheong, Lin Lee
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908343944085504
author Woo, Sangmin
Zhou, Kang
Zhou, Yun
Wang, Shuai
Guan, Sheng
Ding, Haibo
Cheong, Lin Lee
author_facet Woo, Sangmin
Zhou, Kang
Zhou, Yun
Wang, Shuai
Guan, Sheng
Ding, Haibo
Cheong, Lin Lee
contents Large Vision Language Models (LVLMs) often suffer from object hallucination, which undermines their reliability. Surprisingly, we find that simple object-based visual prompting -- overlaying visual cues (e.g., bounding box, circle) on images -- can significantly mitigate such hallucination; however, different visual prompts (VPs) vary in effectiveness. To address this, we propose Black-Box Visual Prompt Engineering (BBVPE), a framework to identify optimal VPs that enhance LVLM responses without needing access to model internals. Our approach employs a pool of candidate VPs and trains a router model to dynamically select the most effective VP for a given input image. This black-box approach is model-agnostic, making it applicable to both open-source and proprietary LVLMs. Evaluations on benchmarks such as POPE and CHAIR demonstrate that BBVPE effectively reduces object hallucination.
format Preprint
id arxiv_https___arxiv_org_abs_2504_21559
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Black-Box Visual Prompt Engineering for Mitigating Object Hallucination in Large Vision Language Models
Woo, Sangmin
Zhou, Kang
Zhou, Yun
Wang, Shuai
Guan, Sheng
Ding, Haibo
Cheong, Lin Lee
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Large Vision Language Models (LVLMs) often suffer from object hallucination, which undermines their reliability. Surprisingly, we find that simple object-based visual prompting -- overlaying visual cues (e.g., bounding box, circle) on images -- can significantly mitigate such hallucination; however, different visual prompts (VPs) vary in effectiveness. To address this, we propose Black-Box Visual Prompt Engineering (BBVPE), a framework to identify optimal VPs that enhance LVLM responses without needing access to model internals. Our approach employs a pool of candidate VPs and trains a router model to dynamically select the most effective VP for a given input image. This black-box approach is model-agnostic, making it applicable to both open-source and proprietary LVLMs. Evaluations on benchmarks such as POPE and CHAIR demonstrate that BBVPE effectively reduces object hallucination.
title Black-Box Visual Prompt Engineering for Mitigating Object Hallucination in Large Vision Language Models
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2504.21559