Distilling Cross-Modal Knowledge via Feature Disentanglement

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Junhong, Zhang, Yuan, Huang, Tao, Xu, Wenchao, Yang, Renyu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912727589453824
author Liu, Junhong
Zhang, Yuan
Huang, Tao
Xu, Wenchao
Yang, Renyu
author_facet Liu, Junhong
Zhang, Yuan
Huang, Tao
Xu, Wenchao
Yang, Renyu
contents Knowledge distillation (KD) has proven highly effective for compressing large models and enhancing the performance of smaller ones. However, its effectiveness diminishes in cross-modal scenarios, such as vision-to-language distillation, where inconsistencies in representation across modalities lead to difficult knowledge transfer. To address this challenge, we propose frequency-decoupled cross-modal knowledge distillation, a method designed to decouple and balance knowledge transfer across modalities by leveraging frequency-domain features. We observed that low-frequency features exhibit high consistency across different modalities, whereas high-frequency features demonstrate extremely low cross-modal similarity. Accordingly, we apply distinct losses to these features: enforcing strong alignment in the low-frequency domain and introducing relaxed alignment for high-frequency features. We also propose a scale consistency loss to address distributional shifts between modalities, and employ a shared classifier to unify feature spaces. Extensive experiments across multiple benchmark datasets show our method substantially outperforms traditional KD and state-of-the-art cross-modal KD approaches. Code is available at https://github.com/Johumliu/FD-CMKD.
format Preprint
id arxiv_https___arxiv_org_abs_2511_19887
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Distilling Cross-Modal Knowledge via Feature Disentanglement
Liu, Junhong
Zhang, Yuan
Huang, Tao
Xu, Wenchao
Yang, Renyu
Computer Vision and Pattern Recognition
Artificial Intelligence
Knowledge distillation (KD) has proven highly effective for compressing large models and enhancing the performance of smaller ones. However, its effectiveness diminishes in cross-modal scenarios, such as vision-to-language distillation, where inconsistencies in representation across modalities lead to difficult knowledge transfer. To address this challenge, we propose frequency-decoupled cross-modal knowledge distillation, a method designed to decouple and balance knowledge transfer across modalities by leveraging frequency-domain features. We observed that low-frequency features exhibit high consistency across different modalities, whereas high-frequency features demonstrate extremely low cross-modal similarity. Accordingly, we apply distinct losses to these features: enforcing strong alignment in the low-frequency domain and introducing relaxed alignment for high-frequency features. We also propose a scale consistency loss to address distributional shifts between modalities, and employ a shared classifier to unify feature spaces. Extensive experiments across multiple benchmark datasets show our method substantially outperforms traditional KD and state-of-the-art cross-modal KD approaches. Code is available at https://github.com/Johumliu/FD-CMKD.
title Distilling Cross-Modal Knowledge via Feature Disentanglement
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2511.19887