MMT-ARD: Multimodal Multi-Teacher Adversarial Distillation for Robust Vision-Language Models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Li, Yuqi, Dong, Junhao, Yang, Chuanguang, Wen, Shiping, Koniusz, Piotr, Huang, Tingwen, Tian, Yingli, Ong, Yew-Soon
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866915631170846720
author Li, Yuqi
Dong, Junhao
Yang, Chuanguang
Wen, Shiping
Koniusz, Piotr
Huang, Tingwen
Tian, Yingli
Ong, Yew-Soon
author_facet Li, Yuqi
Dong, Junhao
Yang, Chuanguang
Wen, Shiping
Koniusz, Piotr
Huang, Tingwen
Tian, Yingli
Ong, Yew-Soon
contents Vision-Language Models (VLMs) are increasingly deployed in safety-critical applications, making their adversarial robustness a crucial concern. While adversarial knowledge distillation has shown promise in transferring robustness from teacher to student models, traditional single-teacher approaches suffer from limited knowledge diversity, slow convergence, and difficulty in balancing robustness and accuracy. To address these challenges, we propose MMT-ARD: a Multimodal Multi-Teacher Adversarial Robust Distillation framework. Our key innovation is a dual-teacher knowledge fusion architecture that collaboratively optimizes clean feature preservation and robust feature enhancement. To better handle challenging adversarial examples, we introduce a dynamic weight allocation strategy based on teacher confidence, enabling adaptive focus on harder samples. Moreover, to mitigate bias among teachers, we design an adaptive sigmoid-based weighting function that balances the strength of knowledge transfer across modalities. Extensive experiments on ImageNet and zero-shot benchmarks demonstrate that MMT-ARD improves robust accuracy by +4.32% and zero-shot accuracy by +3.5% on the ViT-B-32 model, while achieving a 2.3x increase in training efficiency over traditional single-teacher methods. These results highlight the effectiveness and scalability of MMT-ARD in enhancing the adversarial robustness of multimodal large models. Our codes are available at https://github.com/itsnotacie/MMT-ARD.
format Preprint
id arxiv_https___arxiv_org_abs_2511_17448
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MMT-ARD: Multimodal Multi-Teacher Adversarial Distillation for Robust Vision-Language Models
Li, Yuqi
Dong, Junhao
Yang, Chuanguang
Wen, Shiping
Koniusz, Piotr
Huang, Tingwen
Tian, Yingli
Ong, Yew-Soon
Computer Vision and Pattern Recognition
Vision-Language Models (VLMs) are increasingly deployed in safety-critical applications, making their adversarial robustness a crucial concern. While adversarial knowledge distillation has shown promise in transferring robustness from teacher to student models, traditional single-teacher approaches suffer from limited knowledge diversity, slow convergence, and difficulty in balancing robustness and accuracy. To address these challenges, we propose MMT-ARD: a Multimodal Multi-Teacher Adversarial Robust Distillation framework. Our key innovation is a dual-teacher knowledge fusion architecture that collaboratively optimizes clean feature preservation and robust feature enhancement. To better handle challenging adversarial examples, we introduce a dynamic weight allocation strategy based on teacher confidence, enabling adaptive focus on harder samples. Moreover, to mitigate bias among teachers, we design an adaptive sigmoid-based weighting function that balances the strength of knowledge transfer across modalities. Extensive experiments on ImageNet and zero-shot benchmarks demonstrate that MMT-ARD improves robust accuracy by +4.32% and zero-shot accuracy by +3.5% on the ViT-B-32 model, while achieving a 2.3x increase in training efficiency over traditional single-teacher methods. These results highlight the effectiveness and scalability of MMT-ARD in enhancing the adversarial robustness of multimodal large models. Our codes are available at https://github.com/itsnotacie/MMT-ARD.
title MMT-ARD: Multimodal Multi-Teacher Adversarial Distillation for Robust Vision-Language Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.17448