Identifying and Mitigating Position Bias of Multi-image Vision-Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tian, Xinyu, Zou, Shu, Yang, Zhaoyuan, Zhang, Jing
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912280075042816
author Tian, Xinyu
Zou, Shu
Yang, Zhaoyuan
Zhang, Jing
author_facet Tian, Xinyu
Zou, Shu
Yang, Zhaoyuan
Zhang, Jing
contents The evolution of Large Vision-Language Models (LVLMs) has progressed from single to multi-image reasoning. Despite this advancement, our findings indicate that LVLMs struggle to robustly utilize information across multiple images, with predictions significantly affected by the alteration of image positions. To further explore this issue, we introduce Position-wise Question Answering (PQA), a meticulously designed task to quantify reasoning capabilities at each position. Our analysis reveals a pronounced position bias in LVLMs: open-source models excel in reasoning with images positioned later but underperform with those in the middle or at the beginning, while proprietary models show improved comprehension for images at the beginning and end but struggle with those in the middle. Motivated by this, we propose SoFt Attention (SoFA), a simple, training-free approach that mitigates this bias by employing linear interpolation between inter-image causal attention and bidirectional counterparts. Experimental results demonstrate that SoFA reduces position bias and enhances the reasoning performance of existing LVLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2503_13792
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Identifying and Mitigating Position Bias of Multi-image Vision-Language Models
Tian, Xinyu
Zou, Shu
Yang, Zhaoyuan
Zhang, Jing
Computer Vision and Pattern Recognition
The evolution of Large Vision-Language Models (LVLMs) has progressed from single to multi-image reasoning. Despite this advancement, our findings indicate that LVLMs struggle to robustly utilize information across multiple images, with predictions significantly affected by the alteration of image positions. To further explore this issue, we introduce Position-wise Question Answering (PQA), a meticulously designed task to quantify reasoning capabilities at each position. Our analysis reveals a pronounced position bias in LVLMs: open-source models excel in reasoning with images positioned later but underperform with those in the middle or at the beginning, while proprietary models show improved comprehension for images at the beginning and end but struggle with those in the middle. Motivated by this, we propose SoFt Attention (SoFA), a simple, training-free approach that mitigates this bias by employing linear interpolation between inter-image causal attention and bidirectional counterparts. Experimental results demonstrate that SoFA reduces position bias and enhances the reasoning performance of existing LVLMs.
title Identifying and Mitigating Position Bias of Multi-image Vision-Language Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.13792