Exploring Interpretability for Visual Prompt Tuning with Cross-layer Concepts
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866912916679163904 |
|---|---|
| author | Wang, Yubin Jiang, Xinyang Cheng, De Zhao, Xiangqian Wang, Zilong Li, Dongsheng Zhao, Cairong |
| author_facet | Wang, Yubin Jiang, Xinyang Cheng, De Zhao, Xiangqian Wang, Zilong Li, Dongsheng Zhao, Cairong |
| contents | Visual prompt tuning offers significant advantages for adapting pre-trained visual foundation models to specific tasks. However, current research provides limited insight into the interpretability of this approach, which is essential for enhancing AI reliability and enabling AI-driven knowledge discovery. In this paper, rather than learning abstract prompt embeddings, we propose the first framework, named Interpretable Visual Prompt Tuning (IVPT), to explore interpretability for visual prompts by introducing cross-layer concept prototypes. Specifically, visual prompts are linked to human-understandable semantic concepts, represented as a set of category-agnostic prototypes, each corresponding to a specific region of the image. IVPT then aggregates features from these regions to generate interpretable prompts for multiple network layers, allowing the explanation of visual prompts at different network depths and semantic granularities. Comprehensive qualitative and quantitative evaluations on fine-grained classification benchmarks show its superior interpretability and performance over visual prompt tuning methods and existing interpretable methods. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2503_06084 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Exploring Interpretability for Visual Prompt Tuning with Cross-layer Concepts Wang, Yubin Jiang, Xinyang Cheng, De Zhao, Xiangqian Wang, Zilong Li, Dongsheng Zhao, Cairong Computer Vision and Pattern Recognition Visual prompt tuning offers significant advantages for adapting pre-trained visual foundation models to specific tasks. However, current research provides limited insight into the interpretability of this approach, which is essential for enhancing AI reliability and enabling AI-driven knowledge discovery. In this paper, rather than learning abstract prompt embeddings, we propose the first framework, named Interpretable Visual Prompt Tuning (IVPT), to explore interpretability for visual prompts by introducing cross-layer concept prototypes. Specifically, visual prompts are linked to human-understandable semantic concepts, represented as a set of category-agnostic prototypes, each corresponding to a specific region of the image. IVPT then aggregates features from these regions to generate interpretable prompts for multiple network layers, allowing the explanation of visual prompts at different network depths and semantic granularities. Comprehensive qualitative and quantitative evaluations on fine-grained classification benchmarks show its superior interpretability and performance over visual prompt tuning methods and existing interpretable methods. |
| title | Exploring Interpretability for Visual Prompt Tuning with Cross-layer Concepts |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2503.06084 |