Mechanistic Diagnostics of Spatial Lexical Bias in Multimodal Large Language Model Spatial Reasoning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ma, Chuang, Liu, Qianying, Obuchi, Tomoyuki, Cheng, Fei, Yang, Wang, Cai, Sudong, Zheng, Shuyuan, Aizawa, Akiko, Kurohashi, Sadao
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913178280001536
author Ma, Chuang
Liu, Qianying
Obuchi, Tomoyuki
Cheng, Fei
Yang, Wang
Cai, Sudong
Zheng, Shuyuan
Aizawa, Akiko
Kurohashi, Sadao
author_facet Ma, Chuang
Liu, Qianying
Obuchi, Tomoyuki
Cheng, Fei
Yang, Wang
Cai, Sudong
Zheng, Shuyuan
Aizawa, Akiko
Kurohashi, Sadao
contents Multimodal large language models (MLLMs) remain unreliable on spatial multiple-choice questions, and their failures are often attributed to poorly attended visual information. In this work, we identify a complementary failure mode, spatial lexical bias: adding a spatial relation word to the answer options can attract the model's decision and make the newly added option likely to be selected. Using nine open-weight MLLMs, we show that this phenomenon is widely observed. In particular, models can answer a binary spatial question correctly, yet consistently select an incorrect third spatial option once it is added to the answer set. We isolate such binary-stable but ternary-fragile cases as diagnostic examples and leverage mechanistic interpretability tools, revealing that a substantial part of the failure instead originates on the language side rather than the visual side: visual attention analyses and residual-stream probes show the correct spatial relation remains internally available on these failures, while irrelevant-option controls, activation patching, and sparse component interventions trace the bias to specific LLM-side channels and neurons. Based on this finding, we show that a lightweight LLM-only DPO update on tiny single-object-pair synthetic data mitigates the bias, lifting four-way robust accuracy by up to 100 points on synthetic data, and by 68.0, 32.6, and 20.1 points on broader evaluation datasets WhatsUp, SpatialMQA-Direct, and VSR.
format Preprint
id arxiv_https___arxiv_org_abs_2606_01914
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Mechanistic Diagnostics of Spatial Lexical Bias in Multimodal Large Language Model Spatial Reasoning
Ma, Chuang
Liu, Qianying
Obuchi, Tomoyuki
Cheng, Fei
Yang, Wang
Cai, Sudong
Zheng, Shuyuan
Aizawa, Akiko
Kurohashi, Sadao
Computation and Language
Computer Vision and Pattern Recognition
Multimodal large language models (MLLMs) remain unreliable on spatial multiple-choice questions, and their failures are often attributed to poorly attended visual information. In this work, we identify a complementary failure mode, spatial lexical bias: adding a spatial relation word to the answer options can attract the model's decision and make the newly added option likely to be selected. Using nine open-weight MLLMs, we show that this phenomenon is widely observed. In particular, models can answer a binary spatial question correctly, yet consistently select an incorrect third spatial option once it is added to the answer set. We isolate such binary-stable but ternary-fragile cases as diagnostic examples and leverage mechanistic interpretability tools, revealing that a substantial part of the failure instead originates on the language side rather than the visual side: visual attention analyses and residual-stream probes show the correct spatial relation remains internally available on these failures, while irrelevant-option controls, activation patching, and sparse component interventions trace the bias to specific LLM-side channels and neurons. Based on this finding, we show that a lightweight LLM-only DPO update on tiny single-object-pair synthetic data mitigates the bias, lifting four-way robust accuracy by up to 100 points on synthetic data, and by 68.0, 32.6, and 20.1 points on broader evaluation datasets WhatsUp, SpatialMQA-Direct, and VSR.
title Mechanistic Diagnostics of Spatial Lexical Bias in Multimodal Large Language Model Spatial Reasoning
topic Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2606.01914