Investigating Redundancy in Multimodal Large Language Models with Multiple Vision Encoders

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Wang, Yizhou, Mao, Song, Chen, Yang, Shen, Yufan, Yan, Yinqiao, Cai, Pinlong, Wang, Ding, Yan, Guohang, Yu, Zhi, Hu, Xuming, Shi, Botian
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866910020729307136
author Wang, Yizhou
Mao, Song
Chen, Yang
Shen, Yufan
Yan, Yinqiao
Cai, Pinlong
Wang, Ding
Yan, Guohang
Yu, Zhi
Hu, Xuming
Shi, Botian
author_facet Wang, Yizhou
Mao, Song
Chen, Yang
Shen, Yufan
Yan, Yinqiao
Cai, Pinlong
Wang, Ding
Yan, Guohang
Yu, Zhi
Hu, Xuming
Shi, Botian
contents Recent multimodal large language models (MLLMs) increasingly integrate multiple vision encoders to improve performance on various benchmarks, assuming that diverse pretraining objectives yield complementary visual signals. However, we show this assumption often fails in practice. Through systematic encoder masking across representative multi encoder MLLMs, we find that performance typically degrades gracefully, and sometimes even improves, when selected encoders are masked, revealing pervasive encoder redundancy. To quantify this effect, we introduce two principled metrics: the Conditional Utilization Rate (CUR), which measures an encoder s marginal contribution in the presence of others, and the Information Gap (IG), which captures heterogeneity in encoder utility within a model. Using these tools, we observe: (i) strong specialization on tasks like OCR and Chart, where a single encoder can dominate with a CUR greater than 90 percent, (ii) high redundancy on general VQA and knowledge based tasks, where encoders are largely interchangeable, (iii) instances of detrimental encoders with negative CUR. Notably, masking specific encoders can yield up to 16 percent higher accuracy on a specific task category and 3.6 percent overall performance boost compared to the full model.Furthermore, single and dual encoder variants recover over 90 percent of baseline on most non OCR tasks with substantially lower training resources and inference latency. Our analysis challenges the more encoders are better heuristic in MLLMs and provides actionable diagnostics for developing more efficient and effective multimodal architectures.
format Preprint
id arxiv_https___arxiv_org_abs_2507_03262
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Investigating Redundancy in Multimodal Large Language Models with Multiple Vision Encoders
Wang, Yizhou
Mao, Song
Chen, Yang
Shen, Yufan
Yan, Yinqiao
Cai, Pinlong
Wang, Ding
Yan, Guohang
Yu, Zhi
Hu, Xuming
Shi, Botian
Computer Vision and Pattern Recognition
Artificial Intelligence
Recent multimodal large language models (MLLMs) increasingly integrate multiple vision encoders to improve performance on various benchmarks, assuming that diverse pretraining objectives yield complementary visual signals. However, we show this assumption often fails in practice. Through systematic encoder masking across representative multi encoder MLLMs, we find that performance typically degrades gracefully, and sometimes even improves, when selected encoders are masked, revealing pervasive encoder redundancy. To quantify this effect, we introduce two principled metrics: the Conditional Utilization Rate (CUR), which measures an encoder s marginal contribution in the presence of others, and the Information Gap (IG), which captures heterogeneity in encoder utility within a model. Using these tools, we observe: (i) strong specialization on tasks like OCR and Chart, where a single encoder can dominate with a CUR greater than 90 percent, (ii) high redundancy on general VQA and knowledge based tasks, where encoders are largely interchangeable, (iii) instances of detrimental encoders with negative CUR. Notably, masking specific encoders can yield up to 16 percent higher accuracy on a specific task category and 3.6 percent overall performance boost compared to the full model.Furthermore, single and dual encoder variants recover over 90 percent of baseline on most non OCR tasks with substantially lower training resources and inference latency. Our analysis challenges the more encoders are better heuristic in MLLMs and provides actionable diagnostics for developing more efficient and effective multimodal architectures.
title Investigating Redundancy in Multimodal Large Language Models with Multiple Vision Encoders
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2507.03262