AMMKD: Adaptive Multimodal Multi-teacher Distillation for Lightweight Vision-Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Yuqi, Yang, Chuanguang, Dong, Junhao, Yao, Zhengtao, Xu, Haoyan, Dong, Zeyu, Zeng, Hansheng, An, Zhulin, Tian, Yingli |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Frequency-Aligned Knowledge Distillation for Lightweight Spatiotemporal Forecasting
by: Li, Yuqi, et al.
Published: (2025)
by: Li, Yuqi, et al.
Published: (2025)
SRKD: Towards Efficient 3D Point Cloud Segmentation via Structure- and Relation-aware Knowledge Distillation
by: Li, Yuqi, et al.
Published: (2025)
by: Li, Yuqi, et al.
Published: (2025)
MMT-ARD: Multimodal Multi-Teacher Adversarial Distillation for Robust Vision-Language Models
by: Li, Yuqi, et al.
Published: (2025)
by: Li, Yuqi, et al.
Published: (2025)
ChromouVQA: Benchmarking Vision-Language Models under Chromatic Camouflaged Images
by: Zhang, Yunfei, et al.
Published: (2025)
by: Zhang, Yunfei, et al.
Published: (2025)
SGLP: A Similarity Guided Fast Layer Partition Pruning for Compressing Large Deep Models
by: Li, Yuqi, et al.
Published: (2024)
by: Li, Yuqi, et al.
Published: (2024)
Multi-Teacher Knowledge Distillation with Reinforcement Learning for Visual Recognition
by: Yang, Chuanguang, et al.
Published: (2025)
by: Yang, Chuanguang, et al.
Published: (2025)
Prototype-Driven Multi-Feature Generation for Visible-Infrared Person Re-identification
by: Li, Jiarui, et al.
Published: (2024)
by: Li, Jiarui, et al.
Published: (2024)
MPQ-DMv2: Flexible Residual Mixed Precision Quantization for Low-Bit Diffusion Models with Temporal Distillation
by: Feng, Weilun, et al.
Published: (2025)
by: Feng, Weilun, et al.
Published: (2025)
GaitKD: A Universal Decoupled Distillation Framework for Efficient Gait Recognition
by: Li, Yuqi, et al.
Published: (2026)
by: Li, Yuqi, et al.
Published: (2026)
SepPrune: Structured Pruning for Efficient Deep Speech Separation
by: Li, Yuqi, et al.
Published: (2025)
by: Li, Yuqi, et al.
Published: (2025)
Relational Diffusion Distillation for Efficient Image Generation
by: Feng, Weilun, et al.
Published: (2024)
by: Feng, Weilun, et al.
Published: (2024)
S$^2$Q-VDiT: Accurate Quantized Video Diffusion Transformer with Salient Data and Sparse Token Distillation
by: Feng, Weilun, et al.
Published: (2025)
by: Feng, Weilun, et al.
Published: (2025)
Exemplar-Free Class Incremental Learning via Incremental Representation
by: Huang, Libo, et al.
Published: (2024)
by: Huang, Libo, et al.
Published: (2024)
CLIP-KD: An Empirical Study of CLIP Model Distillation
by: Yang, Chuanguang, et al.
Published: (2023)
by: Yang, Chuanguang, et al.
Published: (2023)
MultiAnimate: Pose-Guided Image Animation Made Extensible
by: Hu, Yingcheng, et al.
Published: (2026)
by: Hu, Yingcheng, et al.
Published: (2026)
Distilling Time Series Foundation Models for Efficient Forecasting
by: Li, Yuqi, et al.
Published: (2026)
by: Li, Yuqi, et al.
Published: (2026)
Multi-party Collaborative Attention Control for Image Customization
by: Yang, Han, et al.
Published: (2025)
by: Yang, Han, et al.
Published: (2025)
Teacher-Guided Student Self-Knowledge Distillation Using Diffusion Model
by: Wang, Yu, et al.
Published: (2026)
by: Wang, Yu, et al.
Published: (2026)
GaitProtector: Impersonation-Driven Gait De-Identification via Training-Free Diffusion Latent Optimization
by: Duan, Huiran, et al.
Published: (2026)
by: Duan, Huiran, et al.
Published: (2026)
Efficient Continual Learning through Frequency Decomposition and Integration
by: Liu, Ruiqi, et al.
Published: (2025)
by: Liu, Ruiqi, et al.
Published: (2025)
JEPA-T: Joint-Embedding Predictive Architecture with Text Fusion for Image Generation
by: Wan, Siheng, et al.
Published: (2025)
by: Wan, Siheng, et al.
Published: (2025)
PrePrompt: Predictive prompting for class incremental learning
by: Huang, Libo, et al.
Published: (2025)
by: Huang, Libo, et al.
Published: (2025)
Q-VDiT: Towards Accurate Quantization and Distillation of Video-Generation Diffusion Transformers
by: Feng, Weilun, et al.
Published: (2025)
by: Feng, Weilun, et al.
Published: (2025)
Online Policy Distillation with Decision-Attention
by: Yu, Xinqiang, et al.
Published: (2024)
by: Yu, Xinqiang, et al.
Published: (2024)
Federated Knowledge Distillation for Multi-Model Architectures Lithography Hotspot Detection
by: Li, Yuqi, et al.
Published: (2025)
by: Li, Yuqi, et al.
Published: (2025)
Parameterized Prompt for Incremental Object Detection
by: An, Zijia, et al.
Published: (2025)
by: An, Zijia, et al.
Published: (2025)
DAIT: Distillation from Vision-Language Models to Lightweight Classifiers with Adaptive Intermediate Teacher Transfer
by: He, Zhengxu, et al.
Published: (2026)
by: He, Zhengxu, et al.
Published: (2026)
Echo-Forcing: A Scene Memory Framework for Interactive Long Video Generation
by: Wu, Mingqiang, et al.
Published: (2026)
by: Wu, Mingqiang, et al.
Published: (2026)
QuantSparse: Comprehensively Compressing Video Diffusion Transformer with Model Quantization and Attention Sparsification
by: Feng, Weilun, et al.
Published: (2025)
by: Feng, Weilun, et al.
Published: (2025)
Quantized Visual Geometry Grounded Transformer
by: Feng, Weilun, et al.
Published: (2025)
by: Feng, Weilun, et al.
Published: (2025)
Lightweight Multimodal Adaptation of Vision Language Models for Species Recognition and Habitat Context Interpretation in Drone Thermal Imagery
by: Chen, Hao, et al.
Published: (2026)
by: Chen, Hao, et al.
Published: (2026)
DDTime: Dataset Distillation with Spectral Alignment and Information Bottleneck for Time-Series Forecasting
by: Li, Yuqi, et al.
Published: (2025)
by: Li, Yuqi, et al.
Published: (2025)
Asymmetric Cross-Modal Knowledge Distillation: Bridging Modalities with Weak Semantic Consistency
by: Wei, Riling, et al.
Published: (2025)
by: Wei, Riling, et al.
Published: (2025)
CATP: Contextually Adaptive Token Pruning for Efficient and Enhanced Multimodal In-Context Learning
by: Li, Yanshu, et al.
Published: (2025)
by: Li, Yanshu, et al.
Published: (2025)
Language Models Can See Better: Visual Contrastive Decoding For LLM Multimodal Reasoning
by: Pang, Yuqi, et al.
Published: (2025)
by: Pang, Yuqi, et al.
Published: (2025)
WorldCache: Accelerating World Models for Free via Heterogeneous Token Caching
by: Feng, Weilun, et al.
Published: (2026)
by: Feng, Weilun, et al.
Published: (2026)
Adaptive Residual-Update Steering for Low-Overhead Hallucination Mitigation in Large Vision Language Models
by: Zou, Zhengtao, et al.
Published: (2025)
by: Zou, Zhengtao, et al.
Published: (2025)
Revisiting Cross-Architecture Distillation: Adaptive Dual-Teacher Transfer for Lightweight Video Models
by: Peng, Ying, et al.
Published: (2025)
by: Peng, Ying, et al.
Published: (2025)
VLLFL: A Vision-Language Model Based Lightweight Federated Learning Framework for Smart Agriculture
by: Li, Long, et al.
Published: (2025)
by: Li, Long, et al.
Published: (2025)
LIDAR: Lightweight Adaptive Cue-Aware Fusion Vision Mamba for Multimodal Segmentation of Structural Cracks
by: Liu, Hui, et al.
Published: (2025)
by: Liu, Hui, et al.
Published: (2025)
Similar Items
-
Frequency-Aligned Knowledge Distillation for Lightweight Spatiotemporal Forecasting
by: Li, Yuqi, et al.
Published: (2025) -
SRKD: Towards Efficient 3D Point Cloud Segmentation via Structure- and Relation-aware Knowledge Distillation
by: Li, Yuqi, et al.
Published: (2025) -
MMT-ARD: Multimodal Multi-Teacher Adversarial Distillation for Robust Vision-Language Models
by: Li, Yuqi, et al.
Published: (2025) -
ChromouVQA: Benchmarking Vision-Language Models under Chromatic Camouflaged Images
by: Zhang, Yunfei, et al.
Published: (2025) -
SGLP: A Similarity Guided Fast Layer Partition Pruning for Compressing Large Deep Models
by: Li, Yuqi, et al.
Published: (2024)