Explanation-Driven Counterfactual Testing for Faithfulness in Vision-Language Model Explanations

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ding, Sihao, Vasa, Santosh, Ramadwar, Aditi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915527368114176
author Ding, Sihao
Vasa, Santosh
Ramadwar, Aditi
author_facet Ding, Sihao
Vasa, Santosh
Ramadwar, Aditi
contents Vision-Language Models (VLMs) often produce fluent Natural Language Explanations (NLEs) that sound convincing but may not reflect the causal factors driving predictions. This mismatch of plausibility and faithfulness poses technical and governance risks. We introduce Explanation-Driven Counterfactual Testing (EDCT), a fully automated verification procedure for a target VLM that treats the model's own explanation as a falsifiable hypothesis. Given an image-question pair, EDCT: (1) obtains the model's answer and NLE, (2) parses the NLE into testable visual concepts, (3) generates targeted counterfactual edits via generative inpainting, and (4) computes a Counterfactual Consistency Score (CCS) using LLM-assisted analysis of changes in both answers and explanations. Across 120 curated OK-VQA examples and multiple VLMs, EDCT uncovers substantial faithfulness gaps and provides regulator-aligned audit artifacts indicating when cited concepts fail causal tests.
format Preprint
id arxiv_https___arxiv_org_abs_2510_00047
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Explanation-Driven Counterfactual Testing for Faithfulness in Vision-Language Model Explanations
Ding, Sihao
Vasa, Santosh
Ramadwar, Aditi
Computer Vision and Pattern Recognition
Artificial Intelligence
Vision-Language Models (VLMs) often produce fluent Natural Language Explanations (NLEs) that sound convincing but may not reflect the causal factors driving predictions. This mismatch of plausibility and faithfulness poses technical and governance risks. We introduce Explanation-Driven Counterfactual Testing (EDCT), a fully automated verification procedure for a target VLM that treats the model's own explanation as a falsifiable hypothesis. Given an image-question pair, EDCT: (1) obtains the model's answer and NLE, (2) parses the NLE into testable visual concepts, (3) generates targeted counterfactual edits via generative inpainting, and (4) computes a Counterfactual Consistency Score (CCS) using LLM-assisted analysis of changes in both answers and explanations. Across 120 curated OK-VQA examples and multiple VLMs, EDCT uncovers substantial faithfulness gaps and provides regulator-aligned audit artifacts indicating when cited concepts fail causal tests.
title Explanation-Driven Counterfactual Testing for Faithfulness in Vision-Language Model Explanations
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2510.00047