The More You Tell It, The Less It Sees: Anchoring Bias in Vision-Language Models
Fuente:
Zenodo
Salvato in:
| Autore principale: | |
|---|---|
| Natura: | Recurso digital |
| Pubblicazione: |
Zenodo
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866901570210234368 |
|---|---|
| author | Dubey, Mradul |
| author_facet | Dubey, Mradul |
| contents | <p>Adding structured detection metadata to vision-language model prompts systematically degrades visual reasoning due to anchoring bias, and the delivery channel determines the magnitude. Across seven controlled conditions on a surveillance scene, text-encoded bounding boxes dropped visual reasoning to 53%, visual overlays preserved 69%, and cross-modal ID-mapping collapsed to 47%, despite having smaller text to image token ratio. Plausibly positioned fabricated detections pass unchallenged; the metadata cost on visual perception is monotonic for scene description case. This repository contains all raw prompts, model responses, scoring rubrics, and reproducibility artifacts for the study.</p> |
| format | Recurso digital |
| id | zenodo_https___doi_org_10_5281_zenodo_19557723 |
| institution | Zenodo |
| language | |
| publishDate | 2026 |
| publisher | Zenodo |
| record_format | zenodo |
| spellingShingle | The More You Tell It, The Less It Sees: Anchoring Bias in Vision-Language Models Dubey, Mradul anchoring-bias vision-language-models vlm object-detection yolov8 visual-reasoning prompt-engineering computer-vision <p>Adding structured detection metadata to vision-language model prompts systematically degrades visual reasoning due to anchoring bias, and the delivery channel determines the magnitude. Across seven controlled conditions on a surveillance scene, text-encoded bounding boxes dropped visual reasoning to 53%, visual overlays preserved 69%, and cross-modal ID-mapping collapsed to 47%, despite having smaller text to image token ratio. Plausibly positioned fabricated detections pass unchallenged; the metadata cost on visual perception is monotonic for scene description case. This repository contains all raw prompts, model responses, scoring rubrics, and reproducibility artifacts for the study.</p> |
| title | The More You Tell It, The Less It Sees: Anchoring Bias in Vision-Language Models |
| topic | anchoring-bias vision-language-models vlm object-detection yolov8 visual-reasoning prompt-engineering computer-vision |
| url | https://doi.org/10.5281/zenodo.19557723 |