VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Reilly, Dominick, Govind, Manish Kumar, Xue, Le, Das, Srijan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917016994054144
author Reilly, Dominick
Govind, Manish Kumar
Xue, Le
Das, Srijan
author_facet Reilly, Dominick
Govind, Manish Kumar
Xue, Le
Das, Srijan
contents Large Vision-Language Models (VLMs) excel at general visual reasoning tasks but exhibit sharp performance degradation when applied to novel domains with substantial distribution shifts from pretraining data. Existing domain adaptation approaches finetune different VLM components, but this often results in limited domain-specific feature learning or catastrophic forgetting of prior capabilities. To address these issues, we introduce Vision Contextualized Probing (VisCoP), which augments the VLM's vision encoder with a compact set of learnable visual probes. These probes enable efficient domain-specific adaptation with minimal modification to pretrained parameters. We evaluate VisCoP across three challenging domain adaptation settings-cross-view (exocentric to egocentric), cross-modal (RGB to depth), and cross-task (human understanding to robot control). Experiments show that VisCoP consistently outperforms existing adaptation strategies, achieving superior performance on target domains while effectively retaining source-domain knowledge.
format Preprint
id arxiv_https___arxiv_org_abs_2510_13808
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models
Reilly, Dominick
Govind, Manish Kumar
Xue, Le
Das, Srijan
Computer Vision and Pattern Recognition
Large Vision-Language Models (VLMs) excel at general visual reasoning tasks but exhibit sharp performance degradation when applied to novel domains with substantial distribution shifts from pretraining data. Existing domain adaptation approaches finetune different VLM components, but this often results in limited domain-specific feature learning or catastrophic forgetting of prior capabilities. To address these issues, we introduce Vision Contextualized Probing (VisCoP), which augments the VLM's vision encoder with a compact set of learnable visual probes. These probes enable efficient domain-specific adaptation with minimal modification to pretrained parameters. We evaluate VisCoP across three challenging domain adaptation settings-cross-view (exocentric to egocentric), cross-modal (RGB to depth), and cross-task (human understanding to robot control). Experiments show that VisCoP consistently outperforms existing adaptation strategies, achieving superior performance on target domains while effectively retaining source-domain knowledge.
title VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.13808