The Cost of Language: Centroid Erasure Exposes and Exploits Modal Competition in Multimodal Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Paruchuri, Akshay, Chatterjee, Ishan, Fuchs, Henry, Adeli, Ehsan, Didyk, Piotr
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914477053575168
author Paruchuri, Akshay
Chatterjee, Ishan
Fuchs, Henry
Adeli, Ehsan
Didyk, Piotr
author_facet Paruchuri, Akshay
Chatterjee, Ishan
Fuchs, Henry
Adeli, Ehsan
Didyk, Piotr
contents Multimodal language models systematically underperform on visual perception tasks, yet the structure underlying this failure remains poorly understood. We propose centroid replacement, collapsing each token to its nearest K-means centroid, as a controlled probe for modal dependence. Across seven models spanning three architecture families, erasing text centroid structure costs 4$\times$ more accuracy than erasing visual centroid structure, exposing a universal imbalance where language representations overshadow vision even on tasks that demand visual reasoning. We exploit this asymmetry through text centroid contrastive decoding, recovering up to +16.9% accuracy on individual tasks by contrastively decoding against a text-centroid-erased reference. This intervention varies meaningfully with training approaches: standard fine-tuned models show larger gains (+5.6% on average) than preference-optimized models (+1.5% on average). Our findings suggest that modal competition is structurally localized, correctable at inference time without retraining, and quantifiable as a diagnostic signal to guide future multimodal training.
format Preprint
id arxiv_https___arxiv_org_abs_2604_14363
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle The Cost of Language: Centroid Erasure Exposes and Exploits Modal Competition in Multimodal Language Models
Paruchuri, Akshay
Chatterjee, Ishan
Fuchs, Henry
Adeli, Ehsan
Didyk, Piotr
Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
Multimodal language models systematically underperform on visual perception tasks, yet the structure underlying this failure remains poorly understood. We propose centroid replacement, collapsing each token to its nearest K-means centroid, as a controlled probe for modal dependence. Across seven models spanning three architecture families, erasing text centroid structure costs 4$\times$ more accuracy than erasing visual centroid structure, exposing a universal imbalance where language representations overshadow vision even on tasks that demand visual reasoning. We exploit this asymmetry through text centroid contrastive decoding, recovering up to +16.9% accuracy on individual tasks by contrastively decoding against a text-centroid-erased reference. This intervention varies meaningfully with training approaches: standard fine-tuned models show larger gains (+5.6% on average) than preference-optimized models (+1.5% on average). Our findings suggest that modal competition is structurally localized, correctable at inference time without retraining, and quantifiable as a diagnostic signal to guide future multimodal training.
title The Cost of Language: Centroid Erasure Exposes and Exploits Modal Competition in Multimodal Language Models
topic Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2604.14363