An Empirical Study Into What Matters for Calibrating Vision-Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Tu, Weijie, Deng, Weijian, Campbell, Dylan, Gould, Stephen, Gedeon, Tom |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Closer Look at the Robustness of Contrastive Language-Image Pre-Training (CLIP)
by: Tu, Weijie, et al.
Published: (2024)
by: Tu, Weijie, et al.
Published: (2024)
What Does Softmax Probability Tell Us about Classifiers Ranking Across Diverse Test Conditions?
by: Tu, Weijie, et al.
Published: (2024)
by: Tu, Weijie, et al.
Published: (2024)
Toward a Holistic Evaluation of Robustness in CLIP Models
by: Tu, Weijie, et al.
Published: (2024)
by: Tu, Weijie, et al.
Published: (2024)
Confidence and Dispersity as Signals: Unsupervised Model Evaluation and Ranking
by: Deng, Weijian, et al.
Published: (2025)
by: Deng, Weijian, et al.
Published: (2025)
Ranked from Within: Ranking Large Multimodal Models Without Labels
by: Tu, Weijie, et al.
Published: (2024)
by: Tu, Weijie, et al.
Published: (2024)
3D-GPT: Procedural 3D Modeling with Large Language Models
by: Sun, Chunyi, et al.
Published: (2023)
by: Sun, Chunyi, et al.
Published: (2023)
Calibrating Where It Matters: Constrained Temperature Scaling
by: McKenna, Stephen, et al.
Published: (2024)
by: McKenna, Stephen, et al.
Published: (2024)
Taylor Videos for Action Recognition
by: Wang, Lei, et al.
Published: (2024)
by: Wang, Lei, et al.
Published: (2024)
Online Self-Calibration Against Hallucination in Vision-Language Models
by: Chen, Minghui, et al.
Published: (2026)
by: Chen, Minghui, et al.
Published: (2026)
Calibrated Self-Rewarding Vision Language Models
by: Zhou, Yiyang, et al.
Published: (2024)
by: Zhou, Yiyang, et al.
Published: (2024)
TrackNetV4: Enhancing Fast Sports Object Tracking with Motion Attention Maps
by: Raj, Arjun, et al.
Published: (2024)
by: Raj, Arjun, et al.
Published: (2024)
When Spatial meets Temporal in Action Recognition
by: Chen, Huilin, et al.
Published: (2024)
by: Chen, Huilin, et al.
Published: (2024)
Gems: Group Emotion Profiling Through Multimodal Situational Understanding
by: Kataria, Anubhav, et al.
Published: (2025)
by: Kataria, Anubhav, et al.
Published: (2025)
CSGaze: Context-aware Social Gaze Prediction
by: Madan, Surbhi, et al.
Published: (2025)
by: Madan, Surbhi, et al.
Published: (2025)
AgilePruner: An Empirical Study of Attention and Diversity for Adaptive Visual Token Pruning in Large Vision-Language Models
by: Baek, Changwoo, et al.
Published: (2026)
by: Baek, Changwoo, et al.
Published: (2026)
Quantifying Cross-Modality Memorization in Vision-Language Models
by: Wen, Yuxin, et al.
Published: (2025)
by: Wen, Yuxin, et al.
Published: (2025)
Calibrated and Robust Foundation Models for Vision-Language and Medical Image Tasks Under Distribution Shift
by: Khan, Behraj, et al.
Published: (2025)
by: Khan, Behraj, et al.
Published: (2025)
Handling Imbalanced Pseudolabels for Vision-Language Models with Concept Alignment and Confusion-Aware Calibrated Margin
by: Wang, Yuchen, et al.
Published: (2025)
by: Wang, Yuchen, et al.
Published: (2025)
Motion meets Attention: Video Motion Prompts
by: Chen, Qixiang, et al.
Published: (2024)
by: Chen, Qixiang, et al.
Published: (2024)
Adaptive Multi-head Contrastive Learning
by: Wang, Lei, et al.
Published: (2023)
by: Wang, Lei, et al.
Published: (2023)
How Does the Spatial Distribution of Pre-training Data Affect Geospatial Foundation Models?
by: Purohit, Mirali, et al.
Published: (2025)
by: Purohit, Mirali, et al.
Published: (2025)
Denoising Fisher Training For Neural Implicit Samplers
by: Luo, Weijian, et al.
Published: (2024)
by: Luo, Weijian, et al.
Published: (2024)
D-TPT: Dimensional Entropy Maximization for Calibrating Test-Time Prompt Tuning in Vision-Language Models
by: Han, Jisu, et al.
Published: (2025)
by: Han, Jisu, et al.
Published: (2025)
UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation
by: Wang, Jiayun, et al.
Published: (2026)
by: Wang, Jiayun, et al.
Published: (2026)
What Makes "Good" Distractors for Object Hallucination Evaluation in Large Vision-Language Models?
by: Xie, Ming-Kun, et al.
Published: (2025)
by: Xie, Ming-Kun, et al.
Published: (2025)
Diversity Matters: Revisiting Test-Time Compute in Vision-Language Models
by: Tong, Yijie, et al.
Published: (2026)
by: Tong, Yijie, et al.
Published: (2026)
An Empirical Study of World Model Quantization
by: Fu, Zhongqian, et al.
Published: (2026)
by: Fu, Zhongqian, et al.
Published: (2026)
Overconfidence and Calibration in Medical VQA: Empirical Findings and Hallucination-Aware Mitigation
by: Byun, Ji Young, et al.
Published: (2026)
by: Byun, Ji Young, et al.
Published: (2026)
Pretrain Where? Investigating How Pretraining Data Diversity Impacts Geospatial Foundation Model Performance
by: Kaur, Amandeep, et al.
Published: (2026)
by: Kaur, Amandeep, et al.
Published: (2026)
RefRef: A Synthetic Dataset and Benchmark for Reconstructing Refractive and Reflective Objects
by: Yin, Yue, et al.
Published: (2025)
by: Yin, Yue, et al.
Published: (2025)
Self-Calibrated Tuning of Vision-Language Models for Out-of-Distribution Detection
by: Yu, Geng, et al.
Published: (2024)
by: Yu, Geng, et al.
Published: (2024)
Mint: A Simple Test-Time Adaptation of Vision-Language Models against Common Corruptions
by: Bao, Wenxuan, et al.
Published: (2025)
by: Bao, Wenxuan, et al.
Published: (2025)
Diff-Instruct++: Training One-step Text-to-image Generator Model to Align with Human Preferences
by: Luo, Weijian
Published: (2024)
by: Luo, Weijian
Published: (2024)
Meet JEANIE: a Similarity Measure for 3D Skeleton Sequences via Temporal-Viewpoint Alignment
by: Wang, Lei, et al.
Published: (2024)
by: Wang, Lei, et al.
Published: (2024)
Learnable Expansion of Graph Operators for Multi-Modal Feature Fusion
by: Ding, Dexuan, et al.
Published: (2024)
by: Ding, Dexuan, et al.
Published: (2024)
Routers in Vision Mixture of Experts: An Empirical Study
by: Liu, Tianlin, et al.
Published: (2024)
by: Liu, Tianlin, et al.
Published: (2024)
Beyond Accuracy: What Matters in Designing Well-Behaved Image Classification Models?
by: Hesse, Robin, et al.
Published: (2025)
by: Hesse, Robin, et al.
Published: (2025)
Focusing on What Matters: Object-Agent-centric Tokenization for Vision Language Action models
by: Bendikas, Rokas, et al.
Published: (2025)
by: Bendikas, Rokas, et al.
Published: (2025)
Know What You do Not Know: Verbalized Uncertainty Estimation Robustness on Corrupted Images in Vision-Language Models
by: Borszukovszki, Mirko, et al.
Published: (2025)
by: Borszukovszki, Mirko, et al.
Published: (2025)
Re-Align: Aligning Vision Language Models via Retrieval-Augmented Direct Preference Optimization
by: Xing, Shuo, et al.
Published: (2025)
by: Xing, Shuo, et al.
Published: (2025)
Similar Items
-
A Closer Look at the Robustness of Contrastive Language-Image Pre-Training (CLIP)
by: Tu, Weijie, et al.
Published: (2024) -
What Does Softmax Probability Tell Us about Classifiers Ranking Across Diverse Test Conditions?
by: Tu, Weijie, et al.
Published: (2024) -
Toward a Holistic Evaluation of Robustness in CLIP Models
by: Tu, Weijie, et al.
Published: (2024) -
Confidence and Dispersity as Signals: Unsupervised Model Evaluation and Ranking
by: Deng, Weijian, et al.
Published: (2025) -
Ranked from Within: Ranking Large Multimodal Models Without Labels
by: Tu, Weijie, et al.
Published: (2024)