Instruction Anchor: Dissecting the Mechanistic Dynamics of Modality Arbitration

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Yu, Xu, Mufan, Bai, Xuefeng, Chen, Kehai, Zhang, Pengfei, Xiang, Yang, Zhang, Min
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918491872821248
author Zhang, Yu
Xu, Mufan
Bai, Xuefeng
Chen, Kehai
Zhang, Pengfei
Xiang, Yang
Zhang, Min
author_facet Zhang, Yu
Xu, Mufan
Bai, Xuefeng
Chen, Kehai
Zhang, Pengfei
Xiang, Yang
Zhang, Min
contents Modality following is the ability to selectively leverage multimodal contexts based on user instructions. It is fundamental to the safety and reliability of multimodal large language models (MLLMs) in real-world deployments. However, the internal mechanisms governing this decision-making process remain largely under-explored. In this work, we investigate the mechanism underlying modality following through an information flow perspective. Our findings reveal that instruction tokens serve as structural anchor for modality arbitration: Shallow attention layers perform undifferentiated information transfer, aggregating multimodal cues to instruction tokens as a latent buffer; in contrast, deep attention layers selectively strengthen the instruction-compliant subspace and resolve modality arbitration according to the instruction-specified intent, with a sparse subset of attention heads driving this process. Targeted attention-head interventions further validate the functional specificity of these heads: blocking only $5\%$ of the identified heads substantially degrades modality following while preserving general visual and language capabilities, whereas targeted amplification can restore failed modality-following samples by up to approximately $60\%$. Together, this work provides a mechanistic account of modality following and informs future efforts to improve how MLLMs integrate and utilize multimodal evidence under user instructions.
format Preprint
id arxiv_https___arxiv_org_abs_2602_03677
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Instruction Anchor: Dissecting the Mechanistic Dynamics of Modality Arbitration
Zhang, Yu
Xu, Mufan
Bai, Xuefeng
Chen, Kehai
Zhang, Pengfei
Xiang, Yang
Zhang, Min
Computation and Language
Modality following is the ability to selectively leverage multimodal contexts based on user instructions. It is fundamental to the safety and reliability of multimodal large language models (MLLMs) in real-world deployments. However, the internal mechanisms governing this decision-making process remain largely under-explored. In this work, we investigate the mechanism underlying modality following through an information flow perspective. Our findings reveal that instruction tokens serve as structural anchor for modality arbitration: Shallow attention layers perform undifferentiated information transfer, aggregating multimodal cues to instruction tokens as a latent buffer; in contrast, deep attention layers selectively strengthen the instruction-compliant subspace and resolve modality arbitration according to the instruction-specified intent, with a sparse subset of attention heads driving this process. Targeted attention-head interventions further validate the functional specificity of these heads: blocking only $5\%$ of the identified heads substantially degrades modality following while preserving general visual and language capabilities, whereas targeted amplification can restore failed modality-following samples by up to approximately $60\%$. Together, this work provides a mechanistic account of modality following and informs future efforts to improve how MLLMs integrate and utilize multimodal evidence under user instructions.
title Instruction Anchor: Dissecting the Mechanistic Dynamics of Modality Arbitration
topic Computation and Language
url https://arxiv.org/abs/2602.03677