Just Shift It: Test-Time Prototype Shifting for Zero-Shot Generalization with Vision-Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Sui, Elaine, Wang, Xiaohan, Yeung-Levy, Serena |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Connect, Collapse, Corrupt: Learning Cross-Modal Tasks with Uni-Modal Data
von: Zhang, Yuhui, et al.
Veröffentlicht: (2024)
von: Zhang, Yuhui, et al.
Veröffentlicht: (2024)
Automated Generation of Challenging Multiple-Choice Questions for Vision Language Model Evaluation
von: Zhang, Yuhui, et al.
Veröffentlicht: (2025)
von: Zhang, Yuhui, et al.
Veröffentlicht: (2025)
Test-Time Spectrum-Aware Latent Steering for Zero-Shot Generalization in Vision-Language Models
von: Dafnis, Konstantinos M., et al.
Veröffentlicht: (2025)
von: Dafnis, Konstantinos M., et al.
Veröffentlicht: (2025)
Ask, Pose, Unite: Scaling Data Acquisition for Close Interactions with Vision Language Models
von: Bravo-Sánchez, Laura, et al.
Veröffentlicht: (2024)
von: Bravo-Sánchez, Laura, et al.
Veröffentlicht: (2024)
Generalized Category Discovery under Domain Shifts: From Vision to Vision-Language Models
von: Wang, Hongjun, et al.
Veröffentlicht: (2026)
von: Wang, Hongjun, et al.
Veröffentlicht: (2026)
Viewpoint Textual Inversion: Discovering Scene Representations and 3D View Control in 2D Diffusion Models
von: Burgess, James, et al.
Veröffentlicht: (2023)
von: Burgess, James, et al.
Veröffentlicht: (2023)
Why are Visually-Grounded Language Models Bad at Image Classification?
von: Zhang, Yuhui, et al.
Veröffentlicht: (2024)
von: Zhang, Yuhui, et al.
Veröffentlicht: (2024)
NegVQA: Can Vision Language Models Understand Negation?
von: Zhang, Yuhui, et al.
Veröffentlicht: (2025)
von: Zhang, Yuhui, et al.
Veröffentlicht: (2025)
Diverse Prototypical Ensembles Improve Robustness to Subpopulation Shift
von: To, Minh Nguyen Nhat, et al.
Veröffentlicht: (2025)
von: To, Minh Nguyen Nhat, et al.
Veröffentlicht: (2025)
Label Distribution Shift-Aware Prediction Refinement for Test-Time Adaptation
von: Jang, Minguk, et al.
Veröffentlicht: (2024)
von: Jang, Minguk, et al.
Veröffentlicht: (2024)
A Comprehensive Survey on Test-Time Adaptation under Distribution Shifts
von: Liang, Jian, et al.
Veröffentlicht: (2023)
von: Liang, Jian, et al.
Veröffentlicht: (2023)
Video Action Differencing
von: Burgess, James, et al.
Veröffentlicht: (2025)
von: Burgess, James, et al.
Veröffentlicht: (2025)
AutoCLIP: Auto-tuning Zero-Shot Classifiers for Vision-Language Models
von: Metzen, Jan Hendrik, et al.
Veröffentlicht: (2023)
von: Metzen, Jan Hendrik, et al.
Veröffentlicht: (2023)
Zero-shot Action Localization via the Confidence of Large Vision-Language Models
von: Aklilu, Josiah, et al.
Veröffentlicht: (2024)
von: Aklilu, Josiah, et al.
Veröffentlicht: (2024)
Temporal Preference Optimization for Long-Form Video Understanding
von: Li, Rui, et al.
Veröffentlicht: (2025)
von: Li, Rui, et al.
Veröffentlicht: (2025)
Intriguing Differences Between Zero-Shot and Systematic Evaluations of Vision-Language Transformer Models
von: Salman, Shaeke, et al.
Veröffentlicht: (2024)
von: Salman, Shaeke, et al.
Veröffentlicht: (2024)
Anomaly-Aware Vision-Language Adapters for Zero-Shot Anomaly Detection
von: Aqeel, Muhammad, et al.
Veröffentlicht: (2026)
von: Aqeel, Muhammad, et al.
Veröffentlicht: (2026)
MoETTA: Test-Time Adaptation Under Mixed Distribution Shifts with MoE-LayerNorm
von: Fan, Xiao, et al.
Veröffentlicht: (2025)
von: Fan, Xiao, et al.
Veröffentlicht: (2025)
ChatGPT Encounters Morphing Attack Detection: Zero-Shot MAD with Multi-Modal Large Language Models and General Vision Models
von: Zhang, Haoyu, et al.
Veröffentlicht: (2025)
von: Zhang, Haoyu, et al.
Veröffentlicht: (2025)
VisionTS: Visual Masked Autoencoders Are Free-Lunch Zero-Shot Time Series Forecasters
von: Chen, Mouxiang, et al.
Veröffentlicht: (2024)
von: Chen, Mouxiang, et al.
Veröffentlicht: (2024)
Doubly Debiased Test-Time Prompt Tuning for Vision-Language Models
von: Song, Fei, et al.
Veröffentlicht: (2025)
von: Song, Fei, et al.
Veröffentlicht: (2025)
Visual Language Models as Zero-Shot Deepfake Detectors
von: Pirogov, Viacheslav
Veröffentlicht: (2025)
von: Pirogov, Viacheslav
Veröffentlicht: (2025)
A Shift in Perspective on Causality in Domain Generalization
von: Machlanski, Damian, et al.
Veröffentlicht: (2025)
von: Machlanski, Damian, et al.
Veröffentlicht: (2025)
Can Your Generative Model Detect Out-of-Distribution Covariate Shift?
von: Viviers, Christiaan, et al.
Veröffentlicht: (2024)
von: Viviers, Christiaan, et al.
Veröffentlicht: (2024)
Zero-Shot Visual Reasoning by Vision-Language Models: Benchmarking and Analysis
von: Nagar, Aishik, et al.
Veröffentlicht: (2024)
von: Nagar, Aishik, et al.
Veröffentlicht: (2024)
Closing the Modality Gap for Mixed Modality Search
von: Li, Binxu, et al.
Veröffentlicht: (2025)
von: Li, Binxu, et al.
Veröffentlicht: (2025)
ProtoCLIP: Prototype-Aligned Latent Refinement for Robust Zero-Shot Chest X-Ray Classification
von: Kittler, Florian, et al.
Veröffentlicht: (2026)
von: Kittler, Florian, et al.
Veröffentlicht: (2026)
Zero-Shot Generalization of Vision-Based RL Without Data Augmentation
von: Batra, Sumeet, et al.
Veröffentlicht: (2024)
von: Batra, Sumeet, et al.
Veröffentlicht: (2024)
RadDiff: Describing Differences in Radiology Image Sets with Natural Language
von: Shen, Xiaoxian, et al.
Veröffentlicht: (2026)
von: Shen, Xiaoxian, et al.
Veröffentlicht: (2026)
Configuring Data Augmentations to Reduce Variance Shift in Positional Embedding of Vision Transformers
von: Kim, Bum Jun, et al.
Veröffentlicht: (2024)
von: Kim, Bum Jun, et al.
Veröffentlicht: (2024)
LipShiFT: A Certifiably Robust Shift-based Vision Transformer
von: Menon, Rohan, et al.
Veröffentlicht: (2025)
von: Menon, Rohan, et al.
Veröffentlicht: (2025)
ArtifactLens: Hundreds of Labels Are Enough for Artifact Detection with VLMs
von: Burgess, James, et al.
Veröffentlicht: (2026)
von: Burgess, James, et al.
Veröffentlicht: (2026)
Data or Language Supervision: What Makes CLIP Better than DINO?
von: Liu, Yiming, et al.
Veröffentlicht: (2025)
von: Liu, Yiming, et al.
Veröffentlicht: (2025)
Zero-Shot Action Generalization with Limited Observations
von: Alchihabi, Abdullah, et al.
Veröffentlicht: (2025)
von: Alchihabi, Abdullah, et al.
Veröffentlicht: (2025)
VideoAgent: Long-form Video Understanding with Large Language Model as Agent
von: Wang, Xiaohan, et al.
Veröffentlicht: (2024)
von: Wang, Xiaohan, et al.
Veröffentlicht: (2024)
Spurious Feature Eraser: Stabilizing Test-Time Adaptation for Vision-Language Foundation Model
von: Ma, Huan, et al.
Veröffentlicht: (2024)
von: Ma, Huan, et al.
Veröffentlicht: (2024)
Language-Driven Anchors for Zero-Shot Adversarial Robustness
von: Li, Xiao, et al.
Veröffentlicht: (2023)
von: Li, Xiao, et al.
Veröffentlicht: (2023)
Efficient Few-Shot Learning in Remote Sensing: Fusing Vision and Vision-Language Models
von: Chua, Jia Yun, et al.
Veröffentlicht: (2025)
von: Chua, Jia Yun, et al.
Veröffentlicht: (2025)
Weighted Risk Invariance: Domain Generalization under Invariant Feature Shift
von: Wong, Gina, et al.
Veröffentlicht: (2024)
von: Wong, Gina, et al.
Veröffentlicht: (2024)
Adaptive Concept Bottleneck for Foundation Models Under Distribution Shifts
von: Choi, Jihye, et al.
Veröffentlicht: (2024)
von: Choi, Jihye, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Connect, Collapse, Corrupt: Learning Cross-Modal Tasks with Uni-Modal Data
von: Zhang, Yuhui, et al.
Veröffentlicht: (2024) -
Automated Generation of Challenging Multiple-Choice Questions for Vision Language Model Evaluation
von: Zhang, Yuhui, et al.
Veröffentlicht: (2025) -
Test-Time Spectrum-Aware Latent Steering for Zero-Shot Generalization in Vision-Language Models
von: Dafnis, Konstantinos M., et al.
Veröffentlicht: (2025) -
Ask, Pose, Unite: Scaling Data Acquisition for Close Interactions with Vision Language Models
von: Bravo-Sánchez, Laura, et al.
Veröffentlicht: (2024) -
Generalized Category Discovery under Domain Shifts: From Vision to Vision-Language Models
von: Wang, Hongjun, et al.
Veröffentlicht: (2026)