GIFT: A Framework Towards Global Interpretable Faithful Textual Explanations of Vision Classifiers
Fuente:
arXiv
Saved in:
| Main Authors: | Zablocki, Éloi, Gerard, Valentin, Cardiel, Amaia, Gaussier, Eric, Cord, Matthieu, Valle, Eduardo |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LLM-wrapper: Black-Box Semantic-Aware Adaptation of Vision-Language Models for Referring Expression Comprehension
by: Cardiel, Amaia, et al.
Published: (2024)
by: Cardiel, Amaia, et al.
Published: (2024)
MAD: Motion Appearance Decoupling for efficient Driving World Models
by: Rahimi, Ahmad, et al.
Published: (2026)
by: Rahimi, Ahmad, et al.
Published: (2026)
DRIV-EX: Counterfactual Explanations for Driving LLMs
by: Cardiel, Amaia, et al.
Published: (2026)
by: Cardiel, Amaia, et al.
Published: (2026)
GaussRender: Learning 3D Occupancy with Gaussian Rendering
by: Chambon, Loïck, et al.
Published: (2025)
by: Chambon, Loïck, et al.
Published: (2025)
Annealed Winner-Takes-All for Motion Forecasting
by: Xu, Yihong, et al.
Published: (2024)
by: Xu, Yihong, et al.
Published: (2024)
UniTraj: A Unified Framework for Scalable Vehicle Trajectory Prediction
by: Feng, Lan, et al.
Published: (2024)
by: Feng, Lan, et al.
Published: (2024)
ReGentS: Real-World Safety-Critical Driving Scenario Generation Made Stable
by: Yin, Yuan, et al.
Published: (2024)
by: Yin, Yuan, et al.
Published: (2024)
PointBeV: A Sparse Approach to BeV Predictions
by: Chambon, Loick, et al.
Published: (2023)
by: Chambon, Loick, et al.
Published: (2023)
NAF: Zero-Shot Feature Upsampling via Neighborhood Attention Filtering
by: Chambon, Loick, et al.
Published: (2025)
by: Chambon, Loick, et al.
Published: (2025)
Towards Motion Forecasting with Real-World Perception Inputs: Are End-to-End Approaches Competitive?
by: Xu, Yihong, et al.
Published: (2023)
by: Xu, Yihong, et al.
Published: (2023)
PPT: Pretraining with Pseudo-Labeled Trajectories for Motion Forecasting
by: Xu, Yihong, et al.
Published: (2024)
by: Xu, Yihong, et al.
Published: (2024)
RAP: 3D Rasterization Augmented End-to-End Planning
by: Feng, Lan, et al.
Published: (2025)
by: Feng, Lan, et al.
Published: (2025)
On the Faithfulness of Vision Transformer Explanations
by: Wu, Junyi, et al.
Published: (2024)
by: Wu, Junyi, et al.
Published: (2024)
Zero-Shot Faithful Textual Explanations via Directional-Derivative Influence on Predictions
by: Yamauchi, Toshinori, et al.
Published: (2026)
by: Yamauchi, Toshinori, et al.
Published: (2026)
Token Transformation Matters: Towards Faithful Post-hoc Explanation for Vision Transformer
by: Wu, Junyi, et al.
Published: (2024)
by: Wu, Junyi, et al.
Published: (2024)
Halton Scheduler For Masked Generative Image Transformer
by: Besnier, Victor, et al.
Published: (2025)
by: Besnier, Victor, et al.
Published: (2025)
Valeo4Cast: A Modular Approach to End-to-End Forecasting
by: Xu, Yihong, et al.
Published: (2024)
by: Xu, Yihong, et al.
Published: (2024)
Unsupervised Object Localization in the Era of Self-Supervised ViTs: A Survey
by: Siméoni, Oriane, et al.
Published: (2023)
by: Siméoni, Oriane, et al.
Published: (2023)
Towards Generalizable Trajectory Prediction Using Dual-Level Representation Learning And Adaptive Prompting
by: Messaoud, Kaouther, et al.
Published: (2025)
by: Messaoud, Kaouther, et al.
Published: (2025)
Explanation-Driven Counterfactual Testing for Faithfulness in Vision-Language Model Explanations
by: Ding, Sihao, et al.
Published: (2025)
by: Ding, Sihao, et al.
Published: (2025)
Test-time Contrastive Concepts for Open-world Semantic Segmentation with Vision-Language Models
by: Wysoczańska, Monika, et al.
Published: (2024)
by: Wysoczańska, Monika, et al.
Published: (2024)
Skipping Computations in Multimodal LLMs
by: Shukor, Mustafa, et al.
Published: (2024)
by: Shukor, Mustafa, et al.
Published: (2024)
GIFT: Global Irreplaceability Frame Targeting for Efficient Video Understanding
by: Ma, Junpeng, et al.
Published: (2026)
by: Ma, Junpeng, et al.
Published: (2026)
Faithful Counterfactual Visual Explanations (FCVE)
by: Khan, Bismillah, et al.
Published: (2025)
by: Khan, Bismillah, et al.
Published: (2025)
[Re] Improving Interpretation Faithfulness for Vision Transformers
by: Kurek, Izabela, et al.
Published: (2025)
by: Kurek, Izabela, et al.
Published: (2025)
SSDD: Single-Step Diffusion Decoder for Efficient Image Tokenization
by: Vallaeys, Théophane, et al.
Published: (2025)
by: Vallaeys, Théophane, et al.
Published: (2025)
SAM-Sode: Towards Faithful Explanations for Tiny Bacteria Detection
by: Tan, Wanying, et al.
Published: (2026)
by: Tan, Wanying, et al.
Published: (2026)
Improving Interpretation Faithfulness for Vision Transformers
by: Hu, Lijie, et al.
Published: (2023)
by: Hu, Lijie, et al.
Published: (2023)
VaViM and VaVAM: Autonomous Driving through Video Generative Modeling
by: Bartoccioni, Florent, et al.
Published: (2025)
by: Bartoccioni, Florent, et al.
Published: (2025)
Implicit Multimodal Alignment: On the Generalization of Frozen LLMs to Multimodal Inputs
by: Shukor, Mustafa, et al.
Published: (2024)
by: Shukor, Mustafa, et al.
Published: (2024)
Synthetic Data is an Elegant GIFT for Continual Vision-Language Models
by: Wu, Bin, et al.
Published: (2025)
by: Wu, Bin, et al.
Published: (2025)
Improved Baselines for Data-efficient Perceptual Augmentation of LLMs
by: Vallaeys, Théophane, et al.
Published: (2024)
by: Vallaeys, Théophane, et al.
Published: (2024)
Grad-ECLIP: Gradient-based Visual and Textual Explanations for CLIP
by: Zhao, Chenyang, et al.
Published: (2025)
by: Zhao, Chenyang, et al.
Published: (2025)
Driving on Registers
by: Kirby, Ellington, et al.
Published: (2026)
by: Kirby, Ellington, et al.
Published: (2026)
MLLM-based Textual Explanations for Face Comparison
by: Sony, Redwan, et al.
Published: (2026)
by: Sony, Redwan, et al.
Published: (2026)
Zero-Shot Textual Explanations via Translating Decision-Critical Features
by: Yamauchi, Toshinori, et al.
Published: (2025)
by: Yamauchi, Toshinori, et al.
Published: (2025)
Towards Artwork Explanation in Large-scale Vision Language Models
by: Hayashi, Kazuki, et al.
Published: (2024)
by: Hayashi, Kazuki, et al.
Published: (2024)
Beyond Task Performance: Evaluating and Reducing the Flaws of Large Multimodal Models with In-Context Learning
by: Shukor, Mustafa, et al.
Published: (2023)
by: Shukor, Mustafa, et al.
Published: (2023)
Saliency-R1: Enforcing Interpretable and Faithful Vision-language Reasoning via Saliency-map Alignment Reward
by: Gong, Shizhan, et al.
Published: (2026)
by: Gong, Shizhan, et al.
Published: (2026)
ManiPose: Manifold-Constrained Multi-Hypothesis 3D Human Pose Estimation
by: Rommel, Cédric, et al.
Published: (2023)
by: Rommel, Cédric, et al.
Published: (2023)
Similar Items
-
LLM-wrapper: Black-Box Semantic-Aware Adaptation of Vision-Language Models for Referring Expression Comprehension
by: Cardiel, Amaia, et al.
Published: (2024) -
MAD: Motion Appearance Decoupling for efficient Driving World Models
by: Rahimi, Ahmad, et al.
Published: (2026) -
DRIV-EX: Counterfactual Explanations for Driving LLMs
by: Cardiel, Amaia, et al.
Published: (2026) -
GaussRender: Learning 3D Occupancy with Gaussian Rendering
by: Chambon, Loïck, et al.
Published: (2025) -
Annealed Winner-Takes-All for Motion Forecasting
by: Xu, Yihong, et al.
Published: (2024)