Explainability for Vision Foundation Models: A Survey
Fuente:
arXiv
Saved in:
| Main Authors: | Kazmierczak, Rémi, Berthier, Eloïse, Frehse, Goran, Franchi, Gianni |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CLIP-QDA: An Explainable Concept Bottleneck Model
by: Kazmierczak, Rémi, et al.
Published: (2023)
by: Kazmierczak, Rémi, et al.
Published: (2023)
Enhancing Concept Localization in CLIP-based Concept Bottleneck Models
by: Kazmierczak, Rémi, et al.
Published: (2025)
by: Kazmierczak, Rémi, et al.
Published: (2025)
Benchmarking XAI Explanations with Human-Aligned Evaluations
by: Kazmierczak, Rémi, et al.
Published: (2024)
by: Kazmierczak, Rémi, et al.
Published: (2024)
Concept-Based Mechanistic Interpretability Using Structured Knowledge Graphs
by: Chorna, Sofiia, et al.
Published: (2025)
by: Chorna, Sofiia, et al.
Published: (2025)
Learning to Generate Training Datasets for Robust Semantic Segmentation
by: Hariat, Marwane, et al.
Published: (2023)
by: Hariat, Marwane, et al.
Published: (2023)
Leveraging Visual Signals for Robust Token-Level Uncertainty in Vision-Language Generation
by: Hoche, Joseph, et al.
Published: (2026)
by: Hoche, Joseph, et al.
Published: (2026)
Frustratingly Easy Test-Time Adaptation of Vision-Language Models
by: Farina, Matteo, et al.
Published: (2024)
by: Farina, Matteo, et al.
Published: (2024)
EVLF-FM: Explainable Vision Language Foundation Model for Medicine
by: Bai, Yang, et al.
Published: (2025)
by: Bai, Yang, et al.
Published: (2025)
Hierarchical Light Transformer Ensembles for Multimodal Trajectory Forecasting
by: Lafage, Adrien, et al.
Published: (2024)
by: Lafage, Adrien, et al.
Published: (2024)
SS3D: End2End Self-Supervised 3D from Web Videos
by: Hariat, Marwane, et al.
Published: (2026)
by: Hariat, Marwane, et al.
Published: (2026)
Towards Vision-Language Geo-Foundation Model: A Survey
by: Zhou, Yue, et al.
Published: (2024)
by: Zhou, Yue, et al.
Published: (2024)
A Survey on Remote Sensing Foundation Models: From Vision to Multimodality
by: Huang, Ziyue, et al.
Published: (2025)
by: Huang, Ziyue, et al.
Published: (2025)
Vision Foundation Models in Remote Sensing: A Survey
by: Lu, Siqi, et al.
Published: (2024)
by: Lu, Siqi, et al.
Published: (2024)
A Survey on Segment Anything Model (SAM): Vision Foundation Model Meets Prompt Engineering
by: Zhang, Chaoning, et al.
Published: (2023)
by: Zhang, Chaoning, et al.
Published: (2023)
Large Vision-Language Model Alignment and Misalignment: A Survey Through the Lens of Explainability
by: Shu, Dong, et al.
Published: (2025)
by: Shu, Dong, et al.
Published: (2025)
A Geometric Unification of Concept Learning with Concept Cones
by: Rocchi--Henry, Alexandre, et al.
Published: (2025)
by: Rocchi--Henry, Alexandre, et al.
Published: (2025)
Vision-and-Language Navigation Today and Tomorrow: A Survey in the Era of Foundation Models
by: Zhang, Yue, et al.
Published: (2024)
by: Zhang, Yue, et al.
Published: (2024)
How (Mis)calibrated is Your Federated CLIP and What To Do About It?
by: Singha, Mainak, et al.
Published: (2025)
by: Singha, Mainak, et al.
Published: (2025)
Organizing Unstructured Image Collections using Natural Language
by: Liu, Mingxuan, et al.
Published: (2024)
by: Liu, Mingxuan, et al.
Published: (2024)
Towards Unifying Understanding and Generation in the Era of Vision Foundation Models: A Survey from the Autoregression Perspective
by: Xie, Shenghao, et al.
Published: (2024)
by: Xie, Shenghao, et al.
Published: (2024)
An Explainable Biomedical Foundation Model via Large-Scale Concept-Enhanced Vision-Language Pre-training
by: Nie, Yuxiang, et al.
Published: (2025)
by: Nie, Yuxiang, et al.
Published: (2025)
On the Explainability of Vision-Language Models in Art History
by: Schneider, Stefanie
Published: (2026)
by: Schneider, Stefanie
Published: (2026)
Foundation Models for Video Understanding: A Survey
by: Madan, Neelu, et al.
Published: (2024)
by: Madan, Neelu, et al.
Published: (2024)
Vision-Language Models for Vision Tasks: A Survey
by: Zhang, Jingyi, et al.
Published: (2023)
by: Zhang, Jingyi, et al.
Published: (2023)
Does Object Grounding Really Reduce Hallucination of Large Vision-Language Models?
by: Geigle, Gregor, et al.
Published: (2024)
by: Geigle, Gregor, et al.
Published: (2024)
Improving Semantic Uncertainty Quantification in LVLMs with Semantic Gaussian Processes
by: Hoche, Joseph, et al.
Published: (2025)
by: Hoche, Joseph, et al.
Published: (2025)
Deep Learning for Robust and Explainable Models in Computer Vision
by: Amirian, Mohammadreza
Published: (2024)
by: Amirian, Mohammadreza
Published: (2024)
Are Vision Foundation Models Foundational for Electron Microscopy Image Segmentation?
by: Fuster-Barceló, Caterina, et al.
Published: (2026)
by: Fuster-Barceló, Caterina, et al.
Published: (2026)
Cracks in the Foundation: A Civil Infrastructure Dataset to Challenge Vision Foundation Models
by: Farronato, Nicola, et al.
Published: (2026)
by: Farronato, Nicola, et al.
Published: (2026)
Sapiens: Foundation for Human Vision Models
by: Khirodkar, Rawal, et al.
Published: (2024)
by: Khirodkar, Rawal, et al.
Published: (2024)
Image Segmentation in Foundation Model Era: A Survey
by: Zhou, Tianfei, et al.
Published: (2024)
by: Zhou, Tianfei, et al.
Published: (2024)
Foundation Models for Biomedical Image Segmentation: A Survey
by: Lee, Ho Hin, et al.
Published: (2024)
by: Lee, Ho Hin, et al.
Published: (2024)
African or European Swallow? Benchmarking Large Vision-Language Models for Fine-Grained Object Classification
by: Geigle, Gregor, et al.
Published: (2024)
by: Geigle, Gregor, et al.
Published: (2024)
Vision Superalignment: Weak-to-Strong Generalization for Vision Foundation Models
by: Guo, Jianyuan, et al.
Published: (2024)
by: Guo, Jianyuan, et al.
Published: (2024)
From Local Geometry to Global Pseudo Labeling for Robust Positive Unlabeled Learning under Covariate Shift
by: Gabetni, Firas, et al.
Published: (2026)
by: Gabetni, Firas, et al.
Published: (2026)
SurgXBench: Explainable Vision-Language Model Benchmark for Surgery
by: Cheng, Jiajun, et al.
Published: (2025)
by: Cheng, Jiajun, et al.
Published: (2025)
Comprehensive Attribution: Inherently Explainable Vision Model with Feature Detector
by: Zhang, Xianren, et al.
Published: (2024)
by: Zhang, Xianren, et al.
Published: (2024)
A Survey on Efficient Vision-Language Models
by: Shinde, Gaurav, et al.
Published: (2025)
by: Shinde, Gaurav, et al.
Published: (2025)
Implicit Modeling for Transferability Estimation of Vision Foundation Models
by: Zheng, Yaoyan, et al.
Published: (2025)
by: Zheng, Yaoyan, et al.
Published: (2025)
A Survey on Foundation-Model-Based Industrial Defect Detection
by: Yang, Tianle, et al.
Published: (2025)
by: Yang, Tianle, et al.
Published: (2025)
Similar Items
-
CLIP-QDA: An Explainable Concept Bottleneck Model
by: Kazmierczak, Rémi, et al.
Published: (2023) -
Enhancing Concept Localization in CLIP-based Concept Bottleneck Models
by: Kazmierczak, Rémi, et al.
Published: (2025) -
Benchmarking XAI Explanations with Human-Aligned Evaluations
by: Kazmierczak, Rémi, et al.
Published: (2024) -
Concept-Based Mechanistic Interpretability Using Structured Knowledge Graphs
by: Chorna, Sofiia, et al.
Published: (2025) -
Learning to Generate Training Datasets for Robust Semantic Segmentation
by: Hariat, Marwane, et al.
Published: (2023)