Beyond Next-Token Alignment: Distilling Multimodal Large Language Models via Token Interactions

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Chen, Lin, Zhao, Xiaoke, Ding, Kun, Feng, Weiwei, Miao, Changtao, Wang, Zili, Guo, Wenxuan, Wang, Ying, Zheng, Kaiyuan, Zhang, Bo, Li, Zhe, Xiang, Shiming
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866915788290523136
author Chen, Lin
Zhao, Xiaoke
Ding, Kun
Feng, Weiwei
Miao, Changtao
Wang, Zili
Guo, Wenxuan
Wang, Ying
Zheng, Kaiyuan
Zhang, Bo
Li, Zhe
Xiang, Shiming
author_facet Chen, Lin
Zhao, Xiaoke
Ding, Kun
Feng, Weiwei
Miao, Changtao
Wang, Zili
Guo, Wenxuan
Wang, Ying
Zheng, Kaiyuan
Zhang, Bo
Li, Zhe
Xiang, Shiming
contents Multimodal Large Language Models (MLLMs) demonstrate impressive cross-modal capabilities, yet their substantial size poses significant deployment challenges. Knowledge distillation (KD) is a promising solution for compressing these models, but existing methods primarily rely on static next-token alignment, neglecting the dynamic token interactions, which embed essential capabilities for multimodal understanding and generation. To this end, we introduce Align-TI, a novel KD framework designed from the perspective of Token Interactions. Our approach is motivated by the insight that MLLMs rely on two primary interactions: vision-instruction token interactions to extract relevant visual information, and intra-response token interactions for coherent generation. Accordingly, Align-TI introduces two components: IVA enables the student model to imitate the teacher's instruction-relevant visual information extract capability by aligning on salient visual regions. TPA captures the teacher's dynamic generative logic by aligning the sequential token-to-token transition probabilities. Extensive experiments demonstrate Align-TI's superiority. Notably, our approach achieves $2.6\%$ relative improvement over Vanilla KD, and our distilled Align-TI-2B even outperforms LLaVA-1.5-7B (a much larger MLLM) by $7.0\%$, establishing a new state-of-the-art distillation framework for training parameter-efficient MLLMs. Code is available at https://github.com/lchen1019/Align-TI.
format Preprint
id arxiv_https___arxiv_org_abs_2602_09483
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Beyond Next-Token Alignment: Distilling Multimodal Large Language Models via Token Interactions
Chen, Lin
Zhao, Xiaoke
Ding, Kun
Feng, Weiwei
Miao, Changtao
Wang, Zili
Guo, Wenxuan
Wang, Ying
Zheng, Kaiyuan
Zhang, Bo
Li, Zhe
Xiang, Shiming
Computer Vision and Pattern Recognition
Multimodal Large Language Models (MLLMs) demonstrate impressive cross-modal capabilities, yet their substantial size poses significant deployment challenges. Knowledge distillation (KD) is a promising solution for compressing these models, but existing methods primarily rely on static next-token alignment, neglecting the dynamic token interactions, which embed essential capabilities for multimodal understanding and generation. To this end, we introduce Align-TI, a novel KD framework designed from the perspective of Token Interactions. Our approach is motivated by the insight that MLLMs rely on two primary interactions: vision-instruction token interactions to extract relevant visual information, and intra-response token interactions for coherent generation. Accordingly, Align-TI introduces two components: IVA enables the student model to imitate the teacher's instruction-relevant visual information extract capability by aligning on salient visual regions. TPA captures the teacher's dynamic generative logic by aligning the sequential token-to-token transition probabilities. Extensive experiments demonstrate Align-TI's superiority. Notably, our approach achieves $2.6\%$ relative improvement over Vanilla KD, and our distilled Align-TI-2B even outperforms LLaVA-1.5-7B (a much larger MLLM) by $7.0\%$, establishing a new state-of-the-art distillation framework for training parameter-efficient MLLMs. Code is available at https://github.com/lchen1019/Align-TI.
title Beyond Next-Token Alignment: Distilling Multimodal Large Language Models via Token Interactions
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2602.09483