The More You Tell It, The Less It Sees: Anchoring Bias in Vision-Language Models

Fuente: Zenodo
Salvato in:
Dettagli Bibliografici
Autore principale: Dubey, Mradul
Natura: Recurso digital
Pubblicazione: Zenodo 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866901570210234368
author Dubey, Mradul
author_facet Dubey, Mradul
contents <p>Adding structured detection metadata to vision-language model prompts systematically degrades visual reasoning due to anchoring bias, and the delivery channel determines the magnitude. Across seven controlled conditions on a surveillance scene, text-encoded bounding boxes dropped visual reasoning to 53%, visual overlays preserved 69%, and cross-modal ID-mapping collapsed to 47%, despite having smaller text to image token ratio. Plausibly positioned fabricated detections pass unchallenged; the metadata cost on visual perception is monotonic for scene description case. This repository contains all raw prompts, model responses, scoring rubrics, and reproducibility artifacts for the study.</p>
format Recurso digital
id zenodo_https___doi_org_10_5281_zenodo_19557723
institution Zenodo
language
publishDate 2026
publisher Zenodo
record_format zenodo
spellingShingle The More You Tell It, The Less It Sees: Anchoring Bias in Vision-Language Models
Dubey, Mradul
anchoring-bias
vision-language-models
vlm
object-detection
yolov8
visual-reasoning
prompt-engineering
computer-vision
<p>Adding structured detection metadata to vision-language model prompts systematically degrades visual reasoning due to anchoring bias, and the delivery channel determines the magnitude. Across seven controlled conditions on a surveillance scene, text-encoded bounding boxes dropped visual reasoning to 53%, visual overlays preserved 69%, and cross-modal ID-mapping collapsed to 47%, despite having smaller text to image token ratio. Plausibly positioned fabricated detections pass unchallenged; the metadata cost on visual perception is monotonic for scene description case. This repository contains all raw prompts, model responses, scoring rubrics, and reproducibility artifacts for the study.</p>
title The More You Tell It, The Less It Sees: Anchoring Bias in Vision-Language Models
topic anchoring-bias
vision-language-models
vlm
object-detection
yolov8
visual-reasoning
prompt-engineering
computer-vision
url https://doi.org/10.5281/zenodo.19557723