Understanding the Effects of Distractors on Reasoning Vision-Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bae, Jiyun, Ok, Hyunjong, Mo, Sangwoo, Lee, Jaeho
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913178463502336
author Bae, Jiyun
Ok, Hyunjong
Mo, Sangwoo
Lee, Jaeho
author_facet Bae, Jiyun
Ok, Hyunjong
Mo, Sangwoo
Lee, Jaeho
contents How does irrelevant information (i.e., distractors) affect test-time scaling in vision-language models (VLMs)? Prior work on text-only language models has shown that textual distractors can intensify inverse scaling, causing models to reason longer but less effective reasoning traces. In this work, we investigate whether similar phenomena arise in multimodal settings. We introduce Idis (Images with distractors), a visual question-answering dataset that systematically varies distractors along semantic and numerical dimensions. Our analyses reveal that visual distractors affect reasoning VLMs in a fundamentally different way from textual distractors: although inverse scaling still emerges, visual distractors reduce accuracy without increasing reasoning length. We further show that attribute counts extracted from reasoning traces provide key insights into how distractors interact with reasoning length and accuracy. As a sanity check, we propose a simple prompting strategy that mitigates distractor-driven predictions in reasoning vision-language models.
format Preprint
id arxiv_https___arxiv_org_abs_2511_21397
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Understanding the Effects of Distractors on Reasoning Vision-Language Models
Bae, Jiyun
Ok, Hyunjong
Mo, Sangwoo
Lee, Jaeho
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
How does irrelevant information (i.e., distractors) affect test-time scaling in vision-language models (VLMs)? Prior work on text-only language models has shown that textual distractors can intensify inverse scaling, causing models to reason longer but less effective reasoning traces. In this work, we investigate whether similar phenomena arise in multimodal settings. We introduce Idis (Images with distractors), a visual question-answering dataset that systematically varies distractors along semantic and numerical dimensions. Our analyses reveal that visual distractors affect reasoning VLMs in a fundamentally different way from textual distractors: although inverse scaling still emerges, visual distractors reduce accuracy without increasing reasoning length. We further show that attribute counts extracted from reasoning traces provide key insights into how distractors interact with reasoning length and accuracy. As a sanity check, we propose a simple prompting strategy that mitigates distractor-driven predictions in reasoning vision-language models.
title Understanding the Effects of Distractors on Reasoning Vision-Language Models
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2511.21397