AMMKD: Adaptive Multimodal Multi-teacher Distillation for Lightweight Vision-Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Yuqi, Yang, Chuanguang, Dong, Junhao, Yao, Zhengtao, Xu, Haoyan, Dong, Zeyu, Zeng, Hansheng, An, Zhulin, Tian, Yingli
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912559649521664
author Li, Yuqi
Yang, Chuanguang
Dong, Junhao
Yao, Zhengtao
Xu, Haoyan
Dong, Zeyu
Zeng, Hansheng
An, Zhulin
Tian, Yingli
author_facet Li, Yuqi
Yang, Chuanguang
Dong, Junhao
Yao, Zhengtao
Xu, Haoyan
Dong, Zeyu
Zeng, Hansheng
An, Zhulin
Tian, Yingli
contents The success of large-scale visual language pretraining (VLP) models has driven widespread adoption of image-text retrieval tasks. However, their deployment on mobile devices remains limited due to large model sizes and computational complexity. We propose Adaptive Multi-Modal Multi-Teacher Knowledge Distillation (AMMKD), a novel framework that integrates multi-modal feature fusion, multi-teacher distillation, and adaptive optimization to deliver lightweight yet effective retrieval models. Specifically, our method begins with a feature fusion network that extracts and merges discriminative features from both the image and text modalities. To reduce model parameters and further improve performance, we design a multi-teacher knowledge distillation framework to pre-train two CLIP teacher models. We decouple modalities by pre-computing and storing text features as class vectors via the teacher text encoder to enhance efficiency. To better align teacher and student outputs, we apply KL scatter for probability distribution matching. Finally, we design an adaptive dynamic weighting scheme that treats multi-teacher distillation as a multi-objective optimization problem. By leveraging gradient space diversity, we dynamically adjust the influence of each teacher, reducing conflicts and guiding the student toward more optimal learning directions. Extensive experiments on three benchmark datasets demonstrate that AMMKD achieves superior performance while significantly reducing model complexity, validating its effectiveness and flexibility.
format Preprint
id arxiv_https___arxiv_org_abs_2509_00039
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle AMMKD: Adaptive Multimodal Multi-teacher Distillation for Lightweight Vision-Language Models
Li, Yuqi
Yang, Chuanguang
Dong, Junhao
Yao, Zhengtao
Xu, Haoyan
Dong, Zeyu
Zeng, Hansheng
An, Zhulin
Tian, Yingli
Computer Vision and Pattern Recognition
The success of large-scale visual language pretraining (VLP) models has driven widespread adoption of image-text retrieval tasks. However, their deployment on mobile devices remains limited due to large model sizes and computational complexity. We propose Adaptive Multi-Modal Multi-Teacher Knowledge Distillation (AMMKD), a novel framework that integrates multi-modal feature fusion, multi-teacher distillation, and adaptive optimization to deliver lightweight yet effective retrieval models. Specifically, our method begins with a feature fusion network that extracts and merges discriminative features from both the image and text modalities. To reduce model parameters and further improve performance, we design a multi-teacher knowledge distillation framework to pre-train two CLIP teacher models. We decouple modalities by pre-computing and storing text features as class vectors via the teacher text encoder to enhance efficiency. To better align teacher and student outputs, we apply KL scatter for probability distribution matching. Finally, we design an adaptive dynamic weighting scheme that treats multi-teacher distillation as a multi-objective optimization problem. By leveraging gradient space diversity, we dynamically adjust the influence of each teacher, reducing conflicts and guiding the student toward more optimal learning directions. Extensive experiments on three benchmark datasets demonstrate that AMMKD achieves superior performance while significantly reducing model complexity, validating its effectiveness and flexibility.
title AMMKD: Adaptive Multimodal Multi-teacher Distillation for Lightweight Vision-Language Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.00039