Exploring Interpretability for Visual Prompt Tuning with Cross-layer Concepts

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Wang, Yubin, Jiang, Xinyang, Cheng, De, Zhao, Xiangqian, Wang, Zilong, Li, Dongsheng, Zhao, Cairong
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912916679163904
author Wang, Yubin
Jiang, Xinyang
Cheng, De
Zhao, Xiangqian
Wang, Zilong
Li, Dongsheng
Zhao, Cairong
author_facet Wang, Yubin
Jiang, Xinyang
Cheng, De
Zhao, Xiangqian
Wang, Zilong
Li, Dongsheng
Zhao, Cairong
contents Visual prompt tuning offers significant advantages for adapting pre-trained visual foundation models to specific tasks. However, current research provides limited insight into the interpretability of this approach, which is essential for enhancing AI reliability and enabling AI-driven knowledge discovery. In this paper, rather than learning abstract prompt embeddings, we propose the first framework, named Interpretable Visual Prompt Tuning (IVPT), to explore interpretability for visual prompts by introducing cross-layer concept prototypes. Specifically, visual prompts are linked to human-understandable semantic concepts, represented as a set of category-agnostic prototypes, each corresponding to a specific region of the image. IVPT then aggregates features from these regions to generate interpretable prompts for multiple network layers, allowing the explanation of visual prompts at different network depths and semantic granularities. Comprehensive qualitative and quantitative evaluations on fine-grained classification benchmarks show its superior interpretability and performance over visual prompt tuning methods and existing interpretable methods.
format Preprint
id arxiv_https___arxiv_org_abs_2503_06084
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Exploring Interpretability for Visual Prompt Tuning with Cross-layer Concepts
Wang, Yubin
Jiang, Xinyang
Cheng, De
Zhao, Xiangqian
Wang, Zilong
Li, Dongsheng
Zhao, Cairong
Computer Vision and Pattern Recognition
Visual prompt tuning offers significant advantages for adapting pre-trained visual foundation models to specific tasks. However, current research provides limited insight into the interpretability of this approach, which is essential for enhancing AI reliability and enabling AI-driven knowledge discovery. In this paper, rather than learning abstract prompt embeddings, we propose the first framework, named Interpretable Visual Prompt Tuning (IVPT), to explore interpretability for visual prompts by introducing cross-layer concept prototypes. Specifically, visual prompts are linked to human-understandable semantic concepts, represented as a set of category-agnostic prototypes, each corresponding to a specific region of the image. IVPT then aggregates features from these regions to generate interpretable prompts for multiple network layers, allowing the explanation of visual prompts at different network depths and semantic granularities. Comprehensive qualitative and quantitative evaluations on fine-grained classification benchmarks show its superior interpretability and performance over visual prompt tuning methods and existing interpretable methods.
title Exploring Interpretability for Visual Prompt Tuning with Cross-layer Concepts
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.06084