VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Reilly, Dominick, Govind, Manish Kumar, Xue, Le, Das, Srijan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
From My View to Yours: Ego-to-Exo Transfer in VLMs for Understanding Activities of Daily Living
by: Reilly, Dominick, et al.
Published: (2025)
by: Reilly, Dominick, et al.
Published: (2025)
UniLACT: Depth-Aware RGB Latent Action Learning for Vision-Language-Action Models
by: Govind, Manish Kumar, et al.
Published: (2026)
by: Govind, Manish Kumar, et al.
Published: (2026)
LLAVIDAL: A Large LAnguage VIsion Model for Daily Activities of Living
by: Reilly, Dominick, et al.
Published: (2024)
by: Reilly, Dominick, et al.
Published: (2024)
Fibottention: Inceptive Visual Representation Learning with Diverse Attention Across Heads
by: Rahimian, Ali K., et al.
Published: (2024)
by: Rahimian, Ali K., et al.
Published: (2024)
SKI Models: Skeleton Induced Vision-Language Embeddings for Understanding Activities of Daily Living
by: Sinha, Arkaprava, et al.
Published: (2025)
by: Sinha, Arkaprava, et al.
Published: (2025)
VisOnlyQA: Large Vision Language Models Still Struggle with Visual Perception of Geometric Information
by: Kamoi, Ryo, et al.
Published: (2024)
by: Kamoi, Ryo, et al.
Published: (2024)
DenVisCoM: Dense Vision Correspondence Mamba for Efficient and Real-time Optical Flow and Stereo Estimation
by: Anand, Tushar, et al.
Published: (2026)
by: Anand, Tushar, et al.
Published: (2026)
VisCon-100K: Leveraging Contextual Web Data for Fine-tuning Vision Language Models
by: Kumar, Gokul Karthik, et al.
Published: (2025)
by: Kumar, Gokul Karthik, et al.
Published: (2025)
Source-Free Domain Adaptation with Vision-Language Prior
by: Tang, Song, et al.
Published: (2026)
by: Tang, Song, et al.
Published: (2026)
KHMP: Frequency-Domain Kalman Refinement for High-Fidelity Human Motion Prediction
by: Wu, Wenhan, et al.
Published: (2026)
by: Wu, Wenhan, et al.
Published: (2026)
MS-Temba: Multi-Scale Temporal Mamba for Understanding Long Untrimmed Videos
by: Sinha, Arkaprava, et al.
Published: (2025)
by: Sinha, Arkaprava, et al.
Published: (2025)
Beyond Pixels: Semi-Supervised Semantic Segmentation with a Multi-scale Patch-based Multi-Label Classifier
by: Howlader, Prantik, et al.
Published: (2024)
by: Howlader, Prantik, et al.
Published: (2024)
Fusion-SSAT: Unleashing the Potential of Self-supervised Auxiliary Task by Feature Fusion for Generalized Deepfake Detection
by: Reddy, Shukesh, et al.
Published: (2026)
by: Reddy, Shukesh, et al.
Published: (2026)
Integrating Frequency-Domain Representations with Low-Rank Adaptation in Vision-Language Models
by: Khan, Md Azim, et al.
Published: (2025)
by: Khan, Md Azim, et al.
Published: (2025)
Self-supervised Auxiliary Learning for Texture and Model-based Hybrid Robust and Fair Featuring in Face Analysis
by: Reddy, Shukesh, et al.
Published: (2024)
by: Reddy, Shukesh, et al.
Published: (2024)
Fusion of Domain-Adapted Vision and Language Models for Medical Visual Question Answering
by: Ha, Cuong Nhat, et al.
Published: (2024)
by: Ha, Cuong Nhat, et al.
Published: (2024)
HAViT: Historical Attention Vision Transformer
by: Banik, Swarnendu, et al.
Published: (2026)
by: Banik, Swarnendu, et al.
Published: (2026)
Adversarial Robustness Analysis of Vision-Language Models in Medical Image Segmentation
by: Budathoki, Anjila, et al.
Published: (2025)
by: Budathoki, Anjila, et al.
Published: (2025)
Reasoning or Pattern Matching? Probing Large Vision-Language Models with Visual Puzzles
by: Lymperaiou, Maria, et al.
Published: (2026)
by: Lymperaiou, Maria, et al.
Published: (2026)
BLEnD-Vis: Benchmarking Multimodal Cultural Understanding in Vision Language Models
by: Tan, Bryan Chen Zhengyu, et al.
Published: (2025)
by: Tan, Bryan Chen Zhengyu, et al.
Published: (2025)
VisBias: Measuring Explicit and Implicit Social Biases in Vision Language Models
by: Huang, Jen-tse, et al.
Published: (2025)
by: Huang, Jen-tse, et al.
Published: (2025)
VisRes Bench: On Evaluating the Visual Reasoning Capabilities of VLMs
by: Törtei, Brigitta Malagurski, et al.
Published: (2025)
by: Törtei, Brigitta Malagurski, et al.
Published: (2025)
VisCRA: A Visual Chain Reasoning Attack for Jailbreaking Multimodal Large Language Models
by: Sima, Bingrui, et al.
Published: (2025)
by: Sima, Bingrui, et al.
Published: (2025)
CoDA: Instructive Chain-of-Domain Adaptation with Severity-Aware Visual Prompt Tuning
by: Gong, Ziyang, et al.
Published: (2024)
by: Gong, Ziyang, et al.
Published: (2024)
VisMin: Visual Minimal-Change Understanding
by: Awal, Rabiul, et al.
Published: (2024)
by: Awal, Rabiul, et al.
Published: (2024)
Advancing Reliable Test-Time Adaptation of Vision-Language Models under Visual Variations
by: Liang, Yiwen, et al.
Published: (2025)
by: Liang, Yiwen, et al.
Published: (2025)
Exploring the Benefits of Vision Foundation Models for Unsupervised Domain Adaptation
by: Englert, Brunó B., et al.
Published: (2024)
by: Englert, Brunó B., et al.
Published: (2024)
Introducing Gating and Context into Temporal Action Detection
by: Reka, Aglind, et al.
Published: (2024)
by: Reka, Aglind, et al.
Published: (2024)
DiffSwap++: 3D Latent-Controlled Diffusion for Identity-Preserving Face Swapping
by: Bondurant, Weston, et al.
Published: (2025)
by: Bondurant, Weston, et al.
Published: (2025)
Histopath-C: Towards Realistic Domain Shifts for Histopathology Vision-Language Adaptation
by: Noori, Mehrdad, et al.
Published: (2026)
by: Noori, Mehrdad, et al.
Published: (2026)
Incremental Open-set Domain Adaptation
by: Rakshit, Sayan, et al.
Published: (2024)
by: Rakshit, Sayan, et al.
Published: (2024)
Interactive Video Generation via Domain Adaptation
by: Rawal, Ishaan, et al.
Published: (2025)
by: Rawal, Ishaan, et al.
Published: (2025)
ViLAaD: Enhancing "Attracting and Dispersing'' Source-Free Domain Adaptation with Vision-and-Language Model
by: Tarashima, Shuhei, et al.
Published: (2025)
by: Tarashima, Shuhei, et al.
Published: (2025)
Harnessing Vision-Language Pretrained Models with Temporal-Aware Adaptation for Referring Video Object Segmentation
by: Zhou, Zikun, et al.
Published: (2024)
by: Zhou, Zikun, et al.
Published: (2024)
Structural Graph Probing of Vision-Language Models
by: He, Haoyu, et al.
Published: (2026)
by: He, Haoyu, et al.
Published: (2026)
Triage: Hierarchical Visual Budgeting for Efficient Video Reasoning in Vision-Language Models
by: Wang, Anmin, et al.
Published: (2026)
by: Wang, Anmin, et al.
Published: (2026)
Co-Teaching for Unsupervised Domain Adaptation and Expansion
by: Lin, Hailan, et al.
Published: (2022)
by: Lin, Hailan, et al.
Published: (2022)
VisPhyWorld: Probing Physical Reasoning via Code-Driven Video Reconstruction
by: Liang, Jiarong, et al.
Published: (2026)
by: Liang, Jiarong, et al.
Published: (2026)
Attribution Analysis Meets Model Editing: Advancing Knowledge Correction in Vision Language Models with VisEdit
by: Chen, Qizhou, et al.
Published: (2024)
by: Chen, Qizhou, et al.
Published: (2024)
Domain Adaptation with a Single Vision-Language Embedding
by: Fahes, Mohammad, et al.
Published: (2024)
by: Fahes, Mohammad, et al.
Published: (2024)
Similar Items
-
From My View to Yours: Ego-to-Exo Transfer in VLMs for Understanding Activities of Daily Living
by: Reilly, Dominick, et al.
Published: (2025) -
UniLACT: Depth-Aware RGB Latent Action Learning for Vision-Language-Action Models
by: Govind, Manish Kumar, et al.
Published: (2026) -
LLAVIDAL: A Large LAnguage VIsion Model for Daily Activities of Living
by: Reilly, Dominick, et al.
Published: (2024) -
Fibottention: Inceptive Visual Representation Learning with Diverse Attention Across Heads
by: Rahimian, Ali K., et al.
Published: (2024) -
SKI Models: Skeleton Induced Vision-Language Embeddings for Understanding Activities of Daily Living
by: Sinha, Arkaprava, et al.
Published: (2025)