Capability $\neq$ Interpretability: Human Interpretability of Vision Foundation Models
Fuente:
arXiv
Saved in:
| Main Authors: | Colin, Julien, Goetschalckx, Lore, Oliver, Nuria, Serre, Thomas |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Choosing the right basis for interpretability: Psychophysical comparison between neuron-based and dictionary-based representations
by: Colin, Julien, et al.
Published: (2024)
by: Colin, Julien, et al.
Published: (2024)
Unlocking Feature Visualization for Deeper Networks with MAgnitude Constrained Optimization
by: Fel, Thomas, et al.
Published: (2023)
by: Fel, Thomas, et al.
Published: (2023)
Causality $\neq$ Decodability, and Vice Versa: Lessons from Interpreting Counting ViTs
by: Huang, Lianghuan, et al.
Published: (2025)
by: Huang, Lianghuan, et al.
Published: (2025)
Reading $\neq$ Seeing: Diagnosing and Closing the Typography Gap in Vision-Language Models
by: Zhou, Heng, et al.
Published: (2026)
by: Zhou, Heng, et al.
Published: (2026)
LVLM-Interpret: An Interpretability Tool for Large Vision-Language Models
by: Stan, Gabriela Ben Melech, et al.
Published: (2024)
by: Stan, Gabriela Ben Melech, et al.
Published: (2024)
Scale Alone Does not Improve Mechanistic Interpretability in Vision Models
by: Zimmermann, Roland S., et al.
Published: (2023)
by: Zimmermann, Roland S., et al.
Published: (2023)
A Vision-Language Foundation Model to Enhance Efficiency of Chest X-ray Interpretation
by: Chen, Zhihong, et al.
Published: (2024)
by: Chen, Zhihong, et al.
Published: (2024)
Increasing the Diversity in RGB-to-Thermal Image Translation for Automotive Applications
by: Wang, Kaili, et al.
Published: (2025)
by: Wang, Kaili, et al.
Published: (2025)
SkyMoE: A Vision-Language Foundation Model for Enhancing Geospatial Interpretation with Mixture of Experts
by: Liu, Jiaqi, et al.
Published: (2025)
by: Liu, Jiaqi, et al.
Published: (2025)
Interpretability-Aware Vision Transformer
by: Qiang, Yao, et al.
Published: (2023)
by: Qiang, Yao, et al.
Published: (2023)
A Versatile Foundation Model for AI-enabled Mammogram Interpretation
by: Huang, Fuxiang, et al.
Published: (2025)
by: Huang, Fuxiang, et al.
Published: (2025)
Guided Attention for Interpretable Motion Captioning
by: Radouane, Karim, et al.
Published: (2023)
by: Radouane, Karim, et al.
Published: (2023)
Sapiens: Foundation for Human Vision Models
by: Khirodkar, Rawal, et al.
Published: (2024)
by: Khirodkar, Rawal, et al.
Published: (2024)
OmniMRI: A Unified Vision--Language Foundation Model for Generalist MRI Interpretation
by: He, Xingxin, et al.
Published: (2025)
by: He, Xingxin, et al.
Published: (2025)
MedVersa: A Generalist Foundation Model for Medical Image Interpretation
by: Zhou, Hong-Yu, et al.
Published: (2024)
by: Zhou, Hong-Yu, et al.
Published: (2024)
Interpreting Object-level Foundation Models via Visual Precision Search
by: Chen, Ruoyu, et al.
Published: (2024)
by: Chen, Ruoyu, et al.
Published: (2024)
Artwork Interpretation with Vision Language Models: A Case Study on Emotions and Emotion Symbols
by: Padó, Sebastian, et al.
Published: (2025)
by: Padó, Sebastian, et al.
Published: (2025)
Interpreting Low-level Vision Models with Causal Effect Maps
by: Hu, Jinfan, et al.
Published: (2024)
by: Hu, Jinfan, et al.
Published: (2024)
Multi-Modal Interpretability for Enhanced Localization in Vision-Language Models
by: Imran, Muhammad, et al.
Published: (2025)
by: Imran, Muhammad, et al.
Published: (2025)
VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis
by: Kang, Donggoo, et al.
Published: (2024)
by: Kang, Donggoo, et al.
Published: (2024)
UniBiomed: A Universal Foundation Model for Grounded Biomedical Image Interpretation
by: Wu, Linshan, et al.
Published: (2025)
by: Wu, Linshan, et al.
Published: (2025)
ComFe: An Interpretable Head for Vision Transformers
by: Mannix, Evelyn J., et al.
Published: (2024)
by: Mannix, Evelyn J., et al.
Published: (2024)
Interpretable Zero-Shot Learning with Locally-Aligned Vision-Language Model
by: Chen, Shiming, et al.
Published: (2025)
by: Chen, Shiming, et al.
Published: (2025)
3VL: Using Trees to Improve Vision-Language Models' Interpretability
by: Yellinek, Nir, et al.
Published: (2023)
by: Yellinek, Nir, et al.
Published: (2023)
Epistemic-aware Vision-Language Foundation Model for Fetal Ultrasound Interpretation
by: He, Xiao, et al.
Published: (2025)
by: He, Xiao, et al.
Published: (2025)
Interpretable Failure Detection with Human-Level Concepts
by: Nguyen, Kien X., et al.
Published: (2025)
by: Nguyen, Kien X., et al.
Published: (2025)
Interpretable Vision Transformers in Image Classification via SVDA
by: Arampatzakis, Vasileios, et al.
Published: (2026)
by: Arampatzakis, Vasileios, et al.
Published: (2026)
Interpretable and Testable Vision Features via Sparse Autoencoders
by: Stevens, Samuel, et al.
Published: (2025)
by: Stevens, Samuel, et al.
Published: (2025)
Event-Level Detection of Surgical Instrument Handovers in Videos with Interpretable Vision Models
by: Katsarou, Katerina, et al.
Published: (2026)
by: Katsarou, Katerina, et al.
Published: (2026)
A Reasoning-Enabled Vision-Language Foundation Model for Chest X-ray Interpretation
by: Zhang, Yabin, et al.
Published: (2026)
by: Zhang, Yabin, et al.
Published: (2026)
Assessing the Geolocation Capabilities, Limitations and Societal Risks of Generative Vision-Language Models
by: Grainge, Oliver, et al.
Published: (2025)
by: Grainge, Oliver, et al.
Published: (2025)
Interpretable Debiasing of Vision-Language Models for Social Fairness
by: An, Na Min, et al.
Published: (2026)
by: An, Na Min, et al.
Published: (2026)
Interpreting Neurons in Deep Vision Networks with Language Models
by: Bai, Nicholas, et al.
Published: (2024)
by: Bai, Nicholas, et al.
Published: (2024)
DINO-QPM: Adapting Visual Foundation Models for Globally Interpretable Image Classification
by: Zimmermann, Robert, et al.
Published: (2026)
by: Zimmermann, Robert, et al.
Published: (2026)
Cognitive Alignment At No Cost: Inducing Human Attention Biases For Interpretable Vision Transformers
by: Knights, Ethan
Published: (2026)
by: Knights, Ethan
Published: (2026)
HemBLIP: A Vision-Language Model for Interpretable Leukemia Cell Morphology Analysis
by: van Logtestijn, Julie, et al.
Published: (2026)
by: van Logtestijn, Julie, et al.
Published: (2026)
Towards Concept-based Interpretability of Skin Lesion Diagnosis using Vision-Language Models
by: Patrício, Cristiano, et al.
Published: (2023)
by: Patrício, Cristiano, et al.
Published: (2023)
Leveraging Vision-Language Large Models for Interpretable Video Action Recognition with Semantic Tokenization
by: Peng, Jingwei, et al.
Published: (2025)
by: Peng, Jingwei, et al.
Published: (2025)
Interpretability Transfer from Language to Vision via Sparse Autoencoders
by: Kravets, Alexey, et al.
Published: (2026)
by: Kravets, Alexey, et al.
Published: (2026)
Interpretable Vision Transformers in Monocular Depth Estimation via SVDA
by: Arampatzakis, Vasileios, et al.
Published: (2026)
by: Arampatzakis, Vasileios, et al.
Published: (2026)
Similar Items
-
Choosing the right basis for interpretability: Psychophysical comparison between neuron-based and dictionary-based representations
by: Colin, Julien, et al.
Published: (2024) -
Unlocking Feature Visualization for Deeper Networks with MAgnitude Constrained Optimization
by: Fel, Thomas, et al.
Published: (2023) -
Causality $\neq$ Decodability, and Vice Versa: Lessons from Interpreting Counting ViTs
by: Huang, Lianghuan, et al.
Published: (2025) -
Reading $\neq$ Seeing: Diagnosing and Closing the Typography Gap in Vision-Language Models
by: Zhou, Heng, et al.
Published: (2026) -
LVLM-Interpret: An Interpretability Tool for Large Vision-Language Models
by: Stan, Gabriela Ben Melech, et al.
Published: (2024)