Predictive Regularization Against Visual Representation Degradation in Multimodal Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Enguang, Wang, Qiang, Wu, Yuanchen, Yan, Ke, Yuan, Xinbin, Ding, Shouhong, Liu, Xialei, Cheng, Ming-Ming
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914413056884736
author Wang, Enguang
Wang, Qiang
Wu, Yuanchen
Yan, Ke
Yuan, Xinbin
Ding, Shouhong
Liu, Xialei
Cheng, Ming-Ming
author_facet Wang, Enguang
Wang, Qiang
Wu, Yuanchen
Yan, Ke
Yuan, Xinbin
Ding, Shouhong
Liu, Xialei
Cheng, Ming-Ming
contents While Multimodal Large Language Models (MLLMs) excel at vision-language tasks, the cost of their language-driven training on internal visual foundational competence remains unclear. In this paper, we conduct a detailed diagnostic analysis to unveil a pervasive issue: visual representation degradation in MLLMs. Specifically, we find that compared to the initial visual features, the visual representation in the middle layers of LLM exhibits both a degradation in global function and patch structure. We attribute this phenomenon to a visual sacrifice driven by the singular text-generation objective, where the model compromises its visual fidelity to optimize for answer generation. We argue that a robust MLLM requires both strong cross-modal reasoning and core visual competence, and propose Predictive Regularization (PRe) to force degraded intermediate features to predict initial visual features, thereby maintaining the inherent visual attributes of the MLLM's internal representations. Extensive experiments confirm that mitigating this visual degradation effectively boosts vision-language performance, underscoring the critical importance of fostering robust internal visual representations within MLLMs for comprehensive multimodal understanding.
format Preprint
id arxiv_https___arxiv_org_abs_2603_20808
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Predictive Regularization Against Visual Representation Degradation in Multimodal Large Language Models
Wang, Enguang
Wang, Qiang
Wu, Yuanchen
Yan, Ke
Yuan, Xinbin
Ding, Shouhong
Liu, Xialei
Cheng, Ming-Ming
Computer Vision and Pattern Recognition
Machine Learning
While Multimodal Large Language Models (MLLMs) excel at vision-language tasks, the cost of their language-driven training on internal visual foundational competence remains unclear. In this paper, we conduct a detailed diagnostic analysis to unveil a pervasive issue: visual representation degradation in MLLMs. Specifically, we find that compared to the initial visual features, the visual representation in the middle layers of LLM exhibits both a degradation in global function and patch structure. We attribute this phenomenon to a visual sacrifice driven by the singular text-generation objective, where the model compromises its visual fidelity to optimize for answer generation. We argue that a robust MLLM requires both strong cross-modal reasoning and core visual competence, and propose Predictive Regularization (PRe) to force degraded intermediate features to predict initial visual features, thereby maintaining the inherent visual attributes of the MLLM's internal representations. Extensive experiments confirm that mitigating this visual degradation effectively boosts vision-language performance, underscoring the critical importance of fostering robust internal visual representations within MLLMs for comprehensive multimodal understanding.
title Predictive Regularization Against Visual Representation Degradation in Multimodal Large Language Models
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2603.20808