From Drop-off to Recovery: A Mechanistic Analysis of Segmentation in MLLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Boyong, Kim, Sanghwan, Akata, Zeynep
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911524897947648
author Wu, Boyong
Kim, Sanghwan
Akata, Zeynep
author_facet Wu, Boyong
Kim, Sanghwan
Akata, Zeynep
contents Multimodal Large Language Models (MLLMs) are increasingly applied to pixel-level vision tasks, yet their intrinsic capacity for spatial understanding remains poorly understood. We investigate segmentation capacity through a layerwise linear probing evaluation across the entire MLLM pipeline: vision encoder, adapter, and LLM. We further conduct an intervention based attention knockout analysis to test whether cross-token attention progressively refines visual representations, and an evaluation of bidirectional attention among image tokens on spatial consistency. Our analysis reveals that the adapter introduces a segmentation representation drop-off, but LLM layers progressively recover through attention-mediated refinement, where correctly classified tokens steer misclassified neighbors toward the correct label. At early image token positions, this recovery is bounded by causal attention, which bidirectional attention among image tokens alleviates. These findings provide a mechanistic account of how MLLMs process visual information for segmentation, informing the design of future segmentation-capable models.
format Preprint
id arxiv_https___arxiv_org_abs_2603_17228
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle From Drop-off to Recovery: A Mechanistic Analysis of Segmentation in MLLMs
Wu, Boyong
Kim, Sanghwan
Akata, Zeynep
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Multimodal Large Language Models (MLLMs) are increasingly applied to pixel-level vision tasks, yet their intrinsic capacity for spatial understanding remains poorly understood. We investigate segmentation capacity through a layerwise linear probing evaluation across the entire MLLM pipeline: vision encoder, adapter, and LLM. We further conduct an intervention based attention knockout analysis to test whether cross-token attention progressively refines visual representations, and an evaluation of bidirectional attention among image tokens on spatial consistency. Our analysis reveals that the adapter introduces a segmentation representation drop-off, but LLM layers progressively recover through attention-mediated refinement, where correctly classified tokens steer misclassified neighbors toward the correct label. At early image token positions, this recovery is bounded by causal attention, which bidirectional attention among image tokens alleviates. These findings provide a mechanistic account of how MLLMs process visual information for segmentation, informing the design of future segmentation-capable models.
title From Drop-off to Recovery: A Mechanistic Analysis of Segmentation in MLLMs
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2603.17228