DeepInsert: Early Layer Bypass for Efficient and Performant Multimodal Understanding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Choraria, Moulik, Wu, Xinbo, Bhimaraju, Akhil, Sekhar, Nitesh, Wu, Yue, Zhang, Xu, Singhal, Prateek, Varshney, Lav R.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917219873587200
author Choraria, Moulik
Wu, Xinbo
Bhimaraju, Akhil
Sekhar, Nitesh
Wu, Yue
Zhang, Xu
Singhal, Prateek
Varshney, Lav R.
author_facet Choraria, Moulik
Wu, Xinbo
Bhimaraju, Akhil
Sekhar, Nitesh
Wu, Yue
Zhang, Xu
Singhal, Prateek
Varshney, Lav R.
contents Hyperscaling of data and parameter count in LLMs is yielding diminishing improvement when weighed against training costs, underlining a growing need for more efficient finetuning and inference without sacrificing performance. This is especially so for multimodal language models (MLMs), where the overhead of processing multimodal tokens can limit their practical viability. Parallely, recent work has uncovered implicit cross-modal alignment in the deeper layers of large MLMs, deepening our understanding of how MLMs process and encode information. Motivated by this, and our observation that MLMs naturally defer most cross-modal token interactions to deeper layers of the model, we propose a simple modification. Instead of concatenation with the language prompt at the start, we insert multimodal tokens directly into the middle, allowing them to entirely bypass the early layers. Our results with diverse modalities, (i) LLaVA \& BLIP for vision, (ii) LTU for audio, and (iii) MoLCA for molecular data, and model sizes, starting from 350M to 13B parameters, indicate that our method reduces both training and inference costs, while at least preserving, if not surpassing the performance of existing baselines.
format Preprint
id arxiv_https___arxiv_org_abs_2504_19327
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DeepInsert: Early Layer Bypass for Efficient and Performant Multimodal Understanding
Choraria, Moulik
Wu, Xinbo
Bhimaraju, Akhil
Sekhar, Nitesh
Wu, Yue
Zhang, Xu
Singhal, Prateek
Varshney, Lav R.
Computer Vision and Pattern Recognition
Artificial Intelligence
Hyperscaling of data and parameter count in LLMs is yielding diminishing improvement when weighed against training costs, underlining a growing need for more efficient finetuning and inference without sacrificing performance. This is especially so for multimodal language models (MLMs), where the overhead of processing multimodal tokens can limit their practical viability. Parallely, recent work has uncovered implicit cross-modal alignment in the deeper layers of large MLMs, deepening our understanding of how MLMs process and encode information. Motivated by this, and our observation that MLMs naturally defer most cross-modal token interactions to deeper layers of the model, we propose a simple modification. Instead of concatenation with the language prompt at the start, we insert multimodal tokens directly into the middle, allowing them to entirely bypass the early layers. Our results with diverse modalities, (i) LLaVA \& BLIP for vision, (ii) LTU for audio, and (iii) MoLCA for molecular data, and model sizes, starting from 350M to 13B parameters, indicate that our method reduces both training and inference costs, while at least preserving, if not surpassing the performance of existing baselines.
title DeepInsert: Early Layer Bypass for Efficient and Performant Multimodal Understanding
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2504.19327