Training a Student Expert via Semi-Supervised Foundation Model Distillation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Taghavi, Pardis, Liu, Tian, Li, Renjie, Langari, Reza, Tu, Zhengzhong
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917384057520128
author Taghavi, Pardis
Liu, Tian
Li, Renjie
Langari, Reza
Tu, Zhengzhong
author_facet Taghavi, Pardis
Liu, Tian
Li, Renjie
Langari, Reza
Tu, Zhengzhong
contents Foundation models deliver strong perception but are often too computationally heavy to deploy, and adapting them typically requires costly annotations. We introduce a semi-supervised knowledge distillation (SSKD) framework that compresses pre-trained vision foundation models (VFMs) into compact experts using limited labeled and abundant unlabeled data, and instantiate it for instance segmentation where per-pixel labels are particularly expensive. The framework unfolds in three stages: (1) domain adaptation of the VFM(s) via self-training with contrastive calibration, (2) knowledge transfer through a unified multi-objective loss, and (3) student refinement to mitigate residual pseudo-label bias. Central to our approach is an instance-aware pixel-wise contrastive loss that fuses mask and class scores to extract informative negatives and enforce clear inter-instance margins. By maintaining this contrastive signal across both adaptation and distillation, we align teacher and student embeddings and more effectively leverage unlabeled images. On Cityscapes and ADE20K, our $\approx 11\times$ smaller student improves over its zero-shot VFM teacher(s) by +11.9 and +8.6 AP, surpasses adapted teacher(s) by +3.4 and +1.5 AP, and outperforms state-of-the-art SSKD methods on benchmarks.
format Preprint
id arxiv_https___arxiv_org_abs_2604_03841
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Training a Student Expert via Semi-Supervised Foundation Model Distillation
Taghavi, Pardis
Liu, Tian
Li, Renjie
Langari, Reza
Tu, Zhengzhong
Computer Vision and Pattern Recognition
68T45
I.4.6; I.2.10; I.2.6
Foundation models deliver strong perception but are often too computationally heavy to deploy, and adapting them typically requires costly annotations. We introduce a semi-supervised knowledge distillation (SSKD) framework that compresses pre-trained vision foundation models (VFMs) into compact experts using limited labeled and abundant unlabeled data, and instantiate it for instance segmentation where per-pixel labels are particularly expensive. The framework unfolds in three stages: (1) domain adaptation of the VFM(s) via self-training with contrastive calibration, (2) knowledge transfer through a unified multi-objective loss, and (3) student refinement to mitigate residual pseudo-label bias. Central to our approach is an instance-aware pixel-wise contrastive loss that fuses mask and class scores to extract informative negatives and enforce clear inter-instance margins. By maintaining this contrastive signal across both adaptation and distillation, we align teacher and student embeddings and more effectively leverage unlabeled images. On Cityscapes and ADE20K, our $\approx 11\times$ smaller student improves over its zero-shot VFM teacher(s) by +11.9 and +8.6 AP, surpasses adapted teacher(s) by +3.4 and +1.5 AP, and outperforms state-of-the-art SSKD methods on benchmarks.
title Training a Student Expert via Semi-Supervised Foundation Model Distillation
topic Computer Vision and Pattern Recognition
68T45
I.4.6; I.2.10; I.2.6
url https://arxiv.org/abs/2604.03841