CLIP-RD: Relative Distillation for Efficient CLIP Knowledge Distillation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chung, Jeannie, Jang, Hanna, Yang, Ingyeong, Hwang, Uiwon, Sim, Jaehyeong
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910155802673152
author Chung, Jeannie
Jang, Hanna
Yang, Ingyeong
Hwang, Uiwon
Sim, Jaehyeong
author_facet Chung, Jeannie
Jang, Hanna
Yang, Ingyeong
Hwang, Uiwon
Sim, Jaehyeong
contents CLIP aligns image and text embeddings via contrastive learning and demonstrates strong zero-shot generalization. Its large-scale architecture requires substantial computational and memory resources, motivating the distillation of its capabilities into lightweight student models. However, existing CLIP distillation methods do not explicitly model multi-directional relational dependencies between teacher and student embeddings, limiting the student's ability to preserve the structural relationships encoded by the teacher. To address this, we propose a relational knowledge distillation framework that introduces two novel methods, Vertical Relational Distillation (VRD) and Cross Relational Distillation (XRD). VRD enforces consistency of teacher-student distillation strength across modalities at the distribution level, while XRD imposes bidirectional symmetry on cross-modal teacher-student similarity distributions. By jointly modeling multi-directional relational structures, CLIP-RD promotes faithful alignment of the student embedding geometry with that of the teacher, outperforming existing methods by 0.8%p.
format Preprint
id arxiv_https___arxiv_org_abs_2603_25383
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle CLIP-RD: Relative Distillation for Efficient CLIP Knowledge Distillation
Chung, Jeannie
Jang, Hanna
Yang, Ingyeong
Hwang, Uiwon
Sim, Jaehyeong
Computer Vision and Pattern Recognition
CLIP aligns image and text embeddings via contrastive learning and demonstrates strong zero-shot generalization. Its large-scale architecture requires substantial computational and memory resources, motivating the distillation of its capabilities into lightweight student models. However, existing CLIP distillation methods do not explicitly model multi-directional relational dependencies between teacher and student embeddings, limiting the student's ability to preserve the structural relationships encoded by the teacher. To address this, we propose a relational knowledge distillation framework that introduces two novel methods, Vertical Relational Distillation (VRD) and Cross Relational Distillation (XRD). VRD enforces consistency of teacher-student distillation strength across modalities at the distribution level, while XRD imposes bidirectional symmetry on cross-modal teacher-student similarity distributions. By jointly modeling multi-directional relational structures, CLIP-RD promotes faithful alignment of the student embedding geometry with that of the teacher, outperforming existing methods by 0.8%p.
title CLIP-RD: Relative Distillation for Efficient CLIP Knowledge Distillation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.25383