Online In-Context Distillation for Low-Resource Vision Language Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Kang, Zhiqi, Aljundi, Rahaf, Dorovatas, Vaggelis, Alahari, Karteek
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914452807352320
author Kang, Zhiqi
Aljundi, Rahaf
Dorovatas, Vaggelis
Alahari, Karteek
author_facet Kang, Zhiqi
Aljundi, Rahaf
Dorovatas, Vaggelis
Alahari, Karteek
contents As the field continues its push for ever more resources, this work turns the spotlight on a critical question: how can vision-language models (VLMs) be adapted to thrive in low-resource, budget-constrained settings? While large VLMs offer strong performance, they are impractical to deploy in such settings. Small VLMs, on the other hand, are efficient but typically require costly fine-tuning to close the performance gap with larger models in the deployment domain. Inspired by the in-context learning framework, we propose an online In-Context Distillation (ICD) method, in which a small VLM collaborates with a stronger teacher model at inference time, distilling its knowledge via sparse demonstrations to efficiently bridge the gap between them. Our method is built on an in-depth analysis that identifies the scale and the choice of models for which vision-language ICL is currently feasible, and demonstrates the advantage of ICL over fine-tuning under constrained compute budgets. We enhance our method with a novel cross-modal demonstration selection strategy, teacher test-time scaling to reduce noise, and student uncertainty conditioning to dynamically populate a demonstration pool and minimize teacher queries. Our ICD method significantly boosts the performance of small models (up to 33%) using scarce teacher annotations (as low as 4%), and competes with the teacher's zero-shot performance.
format Preprint
id arxiv_https___arxiv_org_abs_2510_18117
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Online In-Context Distillation for Low-Resource Vision Language Models
Kang, Zhiqi
Aljundi, Rahaf
Dorovatas, Vaggelis
Alahari, Karteek
Computer Vision and Pattern Recognition
As the field continues its push for ever more resources, this work turns the spotlight on a critical question: how can vision-language models (VLMs) be adapted to thrive in low-resource, budget-constrained settings? While large VLMs offer strong performance, they are impractical to deploy in such settings. Small VLMs, on the other hand, are efficient but typically require costly fine-tuning to close the performance gap with larger models in the deployment domain. Inspired by the in-context learning framework, we propose an online In-Context Distillation (ICD) method, in which a small VLM collaborates with a stronger teacher model at inference time, distilling its knowledge via sparse demonstrations to efficiently bridge the gap between them. Our method is built on an in-depth analysis that identifies the scale and the choice of models for which vision-language ICL is currently feasible, and demonstrates the advantage of ICL over fine-tuning under constrained compute budgets. We enhance our method with a novel cross-modal demonstration selection strategy, teacher test-time scaling to reduce noise, and student uncertainty conditioning to dynamically populate a demonstration pool and minimize teacher queries. Our ICD method significantly boosts the performance of small models (up to 33%) using scarce teacher annotations (as low as 4%), and competes with the teacher's zero-shot performance.
title Online In-Context Distillation for Low-Resource Vision Language Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.18117