ComFe: An Interpretable Head for Vision Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | Mannix, Evelyn J., Hodgkinson, Liam, Bondell, Howard |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Preserving Angles Improves Feature Distillation
by: Mannix, Evelyn J., et al.
Published: (2024)
by: Mannix, Evelyn J., et al.
Published: (2024)
A Mixture of Exemplars Approach for Efficient Out-of-Distribution Detection with Foundation Models
by: Mannix, Evelyn, et al.
Published: (2023)
by: Mannix, Evelyn, et al.
Published: (2023)
An interpretable approach to automating the assessment of biofouling in video footage
by: Mannix, Evelyn J., et al.
Published: (2025)
by: Mannix, Evelyn J., et al.
Published: (2025)
GMAR: Gradient-Driven Multi-Head Attention Rollout for Vision Transformer Interpretability
by: Jo, Sehyeong, et al.
Published: (2025)
by: Jo, Sehyeong, et al.
Published: (2025)
Detecting and recognizing characters in Greek papyri with YOLOv8, DeiT and SimCLR
by: Turnbull, Robert, et al.
Published: (2024)
by: Turnbull, Robert, et al.
Published: (2024)
Improving Vision Transformers by Overlapping Heads in Multi-Head Self-Attention
by: Zhang, Tianxiao, et al.
Published: (2024)
by: Zhang, Tianxiao, et al.
Published: (2024)
Interpretability-Aware Vision Transformer
by: Qiang, Yao, et al.
Published: (2023)
by: Qiang, Yao, et al.
Published: (2023)
HEART-VIT: Hessian-Guided Efficient Dynamic Attention and Token Pruning in Vision Transformer
by: Uddin, Mohammad Helal, et al.
Published: (2025)
by: Uddin, Mohammad Helal, et al.
Published: (2025)
SHViT: Single-Head Vision Transformer with Memory Efficient Macro Design
by: Yun, Seokju, et al.
Published: (2024)
by: Yun, Seokju, et al.
Published: (2024)
Interpretable Vision Transformers in Image Classification via SVDA
by: Arampatzakis, Vasileios, et al.
Published: (2026)
by: Arampatzakis, Vasileios, et al.
Published: (2026)
ComAlign: Compositional Alignment in Vision-Language Models
by: Abdollah, Ali, et al.
Published: (2024)
by: Abdollah, Ali, et al.
Published: (2024)
Interpretable Image Classification with Adaptive Prototype-based Vision Transformers
by: Ma, Chiyu, et al.
Published: (2024)
by: Ma, Chiyu, et al.
Published: (2024)
B-cos Alignment for Inherently Interpretable CNNs and Vision Transformers
by: Böhle, Moritz, et al.
Published: (2023)
by: Böhle, Moritz, et al.
Published: (2023)
Interpretable Vision Transformers in Monocular Depth Estimation via SVDA
by: Arampatzakis, Vasileios, et al.
Published: (2026)
by: Arampatzakis, Vasileios, et al.
Published: (2026)
Hierarchical Vision Transformer with Prototypes for Interpretable Medical Image Classification
by: Gallée, Luisa, et al.
Published: (2025)
by: Gallée, Luisa, et al.
Published: (2025)
Sparse but not Simpler: A Multi-Level Interpretability Analysis of Vision Transformers
by: Zhang, Siyu
Published: (2026)
by: Zhang, Siyu
Published: (2026)
Comprehensive Information Bottleneck for Unveiling Universal Attribution to Interpret Vision Transformers
by: Hong, Jung-Ho, et al.
Published: (2025)
by: Hong, Jung-Ho, et al.
Published: (2025)
DeepHistoViT: An Interpretable Vision Transformer Framework for Histopathological Cancer Classification
by: Mosalpuri, Ravi, et al.
Published: (2026)
by: Mosalpuri, Ravi, et al.
Published: (2026)
[Re] Improving Interpretation Faithfulness for Vision Transformers
by: Kurek, Izabela, et al.
Published: (2025)
by: Kurek, Izabela, et al.
Published: (2025)
Your Vision-Language-Action Model Already Has Attention Heads For Path Deviation Detection
by: Jeong, Jaehwan, et al.
Published: (2026)
by: Jeong, Jaehwan, et al.
Published: (2026)
GaussianAvatars: Photorealistic Head Avatars with Rigged 3D Gaussians
by: Qian, Shenhan, et al.
Published: (2023)
by: Qian, Shenhan, et al.
Published: (2023)
VL-OrdinalFormer: Vision Language Guided Ordinal Transformers for Interpretable Knee Osteoarthritis Grading
by: Ullah, Zahid, et al.
Published: (2025)
by: Ullah, Zahid, et al.
Published: (2025)
Calibrated Self-supervised Vision Transformers Improve Intracranial Arterial Calcification Segmentation from Clinical CT Head Scans
by: Jin, Benjamin, et al.
Published: (2025)
by: Jin, Benjamin, et al.
Published: (2025)
MoCom: Motion-based Inter-MAV Visual Communication Using Event Vision and Spiking Neural Networks
by: Nengbo, Zhang, et al.
Published: (2025)
by: Nengbo, Zhang, et al.
Published: (2025)
Debiasing CLIP: Interpreting and Correcting Bias in Attention Heads
by: Yeo, Wei Jie, et al.
Published: (2025)
by: Yeo, Wei Jie, et al.
Published: (2025)
LVLM-Interpret: An Interpretability Tool for Large Vision-Language Models
by: Stan, Gabriela Ben Melech, et al.
Published: (2024)
by: Stan, Gabriela Ben Melech, et al.
Published: (2024)
Capability $\neq$ Interpretability: Human Interpretability of Vision Foundation Models
by: Colin, Julien, et al.
Published: (2026)
by: Colin, Julien, et al.
Published: (2026)
TokenCom: Vision-Language Model for Multimodal and Multitask Token Communications
by: Jiang, Feibo, et al.
Published: (2026)
by: Jiang, Feibo, et al.
Published: (2026)
LeMeViT: Efficient Vision Transformer with Learnable Meta Tokens for Remote Sensing Image Interpretation
by: Jiang, Wentao, et al.
Published: (2024)
by: Jiang, Wentao, et al.
Published: (2024)
Can Cross-Layer Transcoders Replace Vision Transformer Activations? An Interpretable Perspective on Vision
by: Chatzoudis, Gerasimos, et al.
Published: (2026)
by: Chatzoudis, Gerasimos, et al.
Published: (2026)
Can Retrieval Heads See Images? Multimodal Retrieval Heads in Long-Context Vision-Language Models
by: Li, Aaron Branson Cigres, et al.
Published: (2026)
by: Li, Aaron Branson Cigres, et al.
Published: (2026)
Improving Interpretation Faithfulness for Vision Transformers
by: Hu, Lijie, et al.
Published: (2023)
by: Hu, Lijie, et al.
Published: (2023)
Prompt-CAM: Making Vision Transformers Interpretable for Fine-Grained Analysis
by: Chowdhury, Arpita, et al.
Published: (2025)
by: Chowdhury, Arpita, et al.
Published: (2025)
TransAnaNet: Transformer-based Anatomy Change Prediction Network for Head and Neck Cancer Patient Radiotherapy
by: Chen, Meixu, et al.
Published: (2024)
by: Chen, Meixu, et al.
Published: (2024)
Vision Transformers with Natural Language Semantics
by: Kim, Young Kyung, et al.
Published: (2024)
by: Kim, Young Kyung, et al.
Published: (2024)
ComPtr: Towards Diverse Bi-source Dense Prediction Tasks via A Simple yet General Complementary Transformer
by: Pang, Youwei, et al.
Published: (2023)
by: Pang, Youwei, et al.
Published: (2023)
B-Cos Aligned Transformers Learn Human-Interpretable Features
by: Tran, Manuel, et al.
Published: (2024)
by: Tran, Manuel, et al.
Published: (2024)
Denoising Vision Transformers
by: Yang, Jiawei, et al.
Published: (2024)
by: Yang, Jiawei, et al.
Published: (2024)
Retina Vision Transformer (RetinaViT): Introducing Scaled Patches into Vision Transformers
by: Shu, Yuyang, et al.
Published: (2024)
by: Shu, Yuyang, et al.
Published: (2024)
IFViT: Interpretable Fixed-Length Representation for Fingerprint Matching via Vision Transformer
by: Qiu, Yuhang, et al.
Published: (2024)
by: Qiu, Yuhang, et al.
Published: (2024)
Similar Items
-
Preserving Angles Improves Feature Distillation
by: Mannix, Evelyn J., et al.
Published: (2024) -
A Mixture of Exemplars Approach for Efficient Out-of-Distribution Detection with Foundation Models
by: Mannix, Evelyn, et al.
Published: (2023) -
An interpretable approach to automating the assessment of biofouling in video footage
by: Mannix, Evelyn J., et al.
Published: (2025) -
GMAR: Gradient-Driven Multi-Head Attention Rollout for Vision Transformer Interpretability
by: Jo, Sehyeong, et al.
Published: (2025) -
Detecting and recognizing characters in Greek papyri with YOLOv8, DeiT and SimCLR
by: Turnbull, Robert, et al.
Published: (2024)