Towards Long-window Anchoring in Vision-Language Model Distillation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhou, Haoyi, Li, Shuo, Chen, Tianyu, Song, Qi, Gao, Chonghan, Li, Jianxin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908746319396864
author Zhou, Haoyi
Li, Shuo
Chen, Tianyu
Song, Qi
Gao, Chonghan
Li, Jianxin
author_facet Zhou, Haoyi
Li, Shuo
Chen, Tianyu
Song, Qi
Gao, Chonghan
Li, Jianxin
contents While large vision-language models (VLMs) demonstrate strong long-context understanding, their prevalent small branches fail on linguistics-photography alignment for a limited window size. We discover that knowledge distillation improves students' capability as a complement to Rotary Position Embeddings (RoPE) on window sizes (anchored from large models). Building on this insight, we propose LAid, which directly aims at the transfer of long-range attention mechanisms through two complementary components: (1) a progressive distance-weighted attention matching that dynamically emphasizes longer position differences during training, and (2) a learnable RoPE response gain modulation that selectively amplifies position sensitivity where needed. Extensive experiments across multiple model families demonstrate that LAid-distilled models achieve up to 3.2 times longer effective context windows compared to baseline small models, while maintaining or improving performance on standard VL benchmarks. Spectral analysis also suggests that LAid successfully preserves crucial low-frequency attention components that conventional methods fail to transfer. Our work not only provides practical techniques for building more efficient long-context VLMs but also offers theoretical insights into how positional understanding emerges and transfers during distillation.
format Preprint
id arxiv_https___arxiv_org_abs_2512_21576
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Towards Long-window Anchoring in Vision-Language Model Distillation
Zhou, Haoyi
Li, Shuo
Chen, Tianyu
Song, Qi
Gao, Chonghan
Li, Jianxin
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
While large vision-language models (VLMs) demonstrate strong long-context understanding, their prevalent small branches fail on linguistics-photography alignment for a limited window size. We discover that knowledge distillation improves students' capability as a complement to Rotary Position Embeddings (RoPE) on window sizes (anchored from large models). Building on this insight, we propose LAid, which directly aims at the transfer of long-range attention mechanisms through two complementary components: (1) a progressive distance-weighted attention matching that dynamically emphasizes longer position differences during training, and (2) a learnable RoPE response gain modulation that selectively amplifies position sensitivity where needed. Extensive experiments across multiple model families demonstrate that LAid-distilled models achieve up to 3.2 times longer effective context windows compared to baseline small models, while maintaining or improving performance on standard VL benchmarks. Spectral analysis also suggests that LAid successfully preserves crucial low-frequency attention components that conventional methods fail to transfer. Our work not only provides practical techniques for building more efficient long-context VLMs but also offers theoretical insights into how positional understanding emerges and transfers during distillation.
title Towards Long-window Anchoring in Vision-Language Model Distillation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2512.21576