Test-Time Consistency in Vision Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chou, Shih-Han, Chandhok, Shivam, Little, James J., Sigal, Leonid |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MM-R$^3$: On (In-)Consistency of Vision-Language Models (VLMs)
von: Chou, Shih-Han, et al.
Veröffentlicht: (2024)
von: Chou, Shih-Han, et al.
Veröffentlicht: (2024)
Response Wide Shut: Surprising Observations in Basic Vision Language Model Capabilities
von: Chandhok, Shivam, et al.
Veröffentlicht: (2024)
von: Chandhok, Shivam, et al.
Veröffentlicht: (2024)
Implicit and Explicit Commonsense for Multi-sentence Video Captioning
von: Chou, Shih-Han, et al.
Veröffentlicht: (2023)
von: Chou, Shih-Han, et al.
Veröffentlicht: (2023)
SceneGPT: A Language Model for 3D Scene Understanding
von: Chandhok, Shivam
Veröffentlicht: (2024)
von: Chandhok, Shivam
Veröffentlicht: (2024)
Do Vision-Language Foundational models show Robust Visual Perception?
von: Chandhok, Shivam, et al.
Veröffentlicht: (2024)
von: Chandhok, Shivam, et al.
Veröffentlicht: (2024)
Learning What Matters: Prioritized Concept Learning via Relative Error-driven Sample Selection
von: Chandhok, Shivam, et al.
Veröffentlicht: (2025)
von: Chandhok, Shivam, et al.
Veröffentlicht: (2025)
DSeq-JEPA: Discriminative Sequential Joint-Embedding Predictive Architecture
von: He, Xiangteng, et al.
Veröffentlicht: (2025)
von: He, Xiangteng, et al.
Veröffentlicht: (2025)
Framework-agnostic Semantically-aware Global Reasoning for Segmentation
von: Hossain, Mir Rayat Imtiaz, et al.
Veröffentlicht: (2022)
von: Hossain, Mir Rayat Imtiaz, et al.
Veröffentlicht: (2022)
The Power of One: A Single Example is All it Takes for Segmentation in VLMs
von: Hossain, Mir Rayat Imtiaz, et al.
Veröffentlicht: (2025)
von: Hossain, Mir Rayat Imtiaz, et al.
Veröffentlicht: (2025)
Visual Prompting for Generalized Few-shot Segmentation: A Multi-scale Approach
von: Hossain, Mir Rayat Imtiaz, et al.
Veröffentlicht: (2024)
von: Hossain, Mir Rayat Imtiaz, et al.
Veröffentlicht: (2024)
Tinted Frames: Question Framing Blinds Vision-Language Models
von: Fan, Wan-Cyuan, et al.
Veröffentlicht: (2026)
von: Fan, Wan-Cyuan, et al.
Veröffentlicht: (2026)
Emergent Open-Vocabulary Semantic Segmentation from Off-the-shelf Vision-Language Models
von: Luo, Jiayun, et al.
Veröffentlicht: (2023)
von: Luo, Jiayun, et al.
Veröffentlicht: (2023)
Negation-Aware Test-Time Adaptation for Vision-Language Models
von: Han, Haochen, et al.
Veröffentlicht: (2025)
von: Han, Haochen, et al.
Veröffentlicht: (2025)
MMFactory: A Universal Solution Search Engine for Vision-Language Tasks
von: Fan, Wan-Cyuan, et al.
Veröffentlicht: (2024)
von: Fan, Wan-Cyuan, et al.
Veröffentlicht: (2024)
SCALE-VLP: Soft-Weighted Contrastive Volumetric Vision-Language Pre-training with Spatial-Knowledge Semantics
von: Mahdizadeh, Ailar, et al.
Veröffentlicht: (2025)
von: Mahdizadeh, Ailar, et al.
Veröffentlicht: (2025)
Preventing Catastrophic Forgetting through Memory Networks in Continuous Detection
von: Bhatt, Gaurav, et al.
Veröffentlicht: (2024)
von: Bhatt, Gaurav, et al.
Veröffentlicht: (2024)
TTRV: Test-Time Reinforcement Learning for Vision Language Models
von: Singh, Akshit, et al.
Veröffentlicht: (2025)
von: Singh, Akshit, et al.
Veröffentlicht: (2025)
Factorized Video Autoencoders for Efficient Generative Modelling
von: Suhail, Mohammed, et al.
Veröffentlicht: (2024)
von: Suhail, Mohammed, et al.
Veröffentlicht: (2024)
Realistic Test-Time Adaptation of Vision-Language Models
von: Zanella, Maxime, et al.
Veröffentlicht: (2025)
von: Zanella, Maxime, et al.
Veröffentlicht: (2025)
Bayesian Test-Time Adaptation for Vision-Language Models
von: Zhou, Lihua, et al.
Veröffentlicht: (2025)
von: Zhou, Lihua, et al.
Veröffentlicht: (2025)
Efficient Test-Time Adaptation of Vision-Language Models
von: Karmanov, Adilbek, et al.
Veröffentlicht: (2024)
von: Karmanov, Adilbek, et al.
Veröffentlicht: (2024)
To Sink or Not to Sink: Visual Information Pathways in Large Vision-Language Models
von: Luo, Jiayun, et al.
Veröffentlicht: (2025)
von: Luo, Jiayun, et al.
Veröffentlicht: (2025)
Multi-modal News Understanding with Professionally Labelled Videos (ReutersViLNews)
von: Chou, Shih-Han, et al.
Veröffentlicht: (2024)
von: Chou, Shih-Han, et al.
Veröffentlicht: (2024)
Joint Generative Modeling of Grounded Scene Graphs and Images via Diffusion Models
von: Xu, Bicheng, et al.
Veröffentlicht: (2024)
von: Xu, Bicheng, et al.
Veröffentlicht: (2024)
Ultra-Light Test-Time Adaptation for Vision--Language Models
von: Kim, Byunghyun
Veröffentlicht: (2025)
von: Kim, Byunghyun
Veröffentlicht: (2025)
Online Gaussian Test-Time Adaptation of Vision-Language Models
von: Fuchs, Clément, et al.
Veröffentlicht: (2025)
von: Fuchs, Clément, et al.
Veröffentlicht: (2025)
Flatness Guided Test-Time Adaptation for Vision-Language Models
von: Li, Aodi, et al.
Veröffentlicht: (2025)
von: Li, Aodi, et al.
Veröffentlicht: (2025)
Test-Time Hinting for Black-Box Vision-Language Models
von: Hou, Kaihua, et al.
Veröffentlicht: (2026)
von: Hou, Kaihua, et al.
Veröffentlicht: (2026)
Prototype-Based Test-Time Adaptation of Vision-Language Models
von: Huang, Zhaohong, et al.
Veröffentlicht: (2026)
von: Huang, Zhaohong, et al.
Veröffentlicht: (2026)
Efficient Test-Time Prompt Tuning for Vision-Language Models
von: Zhu, Yuhan, et al.
Veröffentlicht: (2024)
von: Zhu, Yuhan, et al.
Veröffentlicht: (2024)
Advancing Reliable Test-Time Adaptation of Vision-Language Models under Visual Variations
von: Liang, Yiwen, et al.
Veröffentlicht: (2025)
von: Liang, Yiwen, et al.
Veröffentlicht: (2025)
InvAD: Inversion-based Reconstruction-Free Anomaly Detection with Diffusion Models
von: Sakai, Shunsuke, et al.
Veröffentlicht: (2025)
von: Sakai, Shunsuke, et al.
Veröffentlicht: (2025)
Adaptive Cache Enhancement for Test-Time Adaptation of Vision-Language Models
von: Nguyen, Khanh-Binh, et al.
Veröffentlicht: (2025)
von: Nguyen, Khanh-Binh, et al.
Veröffentlicht: (2025)
3VL: Using Trees to Improve Vision-Language Models' Interpretability
von: Yellinek, Nir, et al.
Veröffentlicht: (2023)
von: Yellinek, Nir, et al.
Veröffentlicht: (2023)
All in One: A Unified Synthetic Data Pipeline for Multimodal Video Understanding
von: Rahman, Tanzila, et al.
Veröffentlicht: (2026)
von: Rahman, Tanzila, et al.
Veröffentlicht: (2026)
Consistency-guided Prompt Learning for Vision-Language Models
von: Roy, Shuvendu, et al.
Veröffentlicht: (2023)
von: Roy, Shuvendu, et al.
Veröffentlicht: (2023)
Unveiling the Tapestry of Consistency in Large Vision-Language Models
von: Zhang, Yuan, et al.
Veröffentlicht: (2024)
von: Zhang, Yuan, et al.
Veröffentlicht: (2024)
A Lost Opportunity for Vision-Language Models: A Comparative Study of Online Test-Time Adaptation for Vision-Language Models
von: Döbler, Mario, et al.
Veröffentlicht: (2024)
von: Döbler, Mario, et al.
Veröffentlicht: (2024)
Test-Time Adaptation of Vision-Language Models for Open-Vocabulary Semantic Segmentation
von: Noori, Mehrdad, et al.
Veröffentlicht: (2025)
von: Noori, Mehrdad, et al.
Veröffentlicht: (2025)
Mitigating Cache Noise in Test-Time Adaptation for Large Vision-Language Models
von: Zhai, Haotian, et al.
Veröffentlicht: (2025)
von: Zhai, Haotian, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
MM-R$^3$: On (In-)Consistency of Vision-Language Models (VLMs)
von: Chou, Shih-Han, et al.
Veröffentlicht: (2024) -
Response Wide Shut: Surprising Observations in Basic Vision Language Model Capabilities
von: Chandhok, Shivam, et al.
Veröffentlicht: (2024) -
Implicit and Explicit Commonsense for Multi-sentence Video Captioning
von: Chou, Shih-Han, et al.
Veröffentlicht: (2023) -
SceneGPT: A Language Model for 3D Scene Understanding
von: Chandhok, Shivam
Veröffentlicht: (2024) -
Do Vision-Language Foundational models show Robust Visual Perception?
von: Chandhok, Shivam, et al.
Veröffentlicht: (2024)