DecomCAM: Advancing Beyond Saliency Maps through Decomposition and Integration
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866914815609405440 |
|---|---|
| author | Yang, Yuguang Guo, Runtang Wu, Sheng Wang, Yimi Yang, Linlin Fan, Bo Zhong, Jilong Zhang, Juan Zhang, Baochang |
| author_facet | Yang, Yuguang Guo, Runtang Wu, Sheng Wang, Yimi Yang, Linlin Fan, Bo Zhong, Jilong Zhang, Juan Zhang, Baochang |
| contents | Interpreting complex deep networks, notably pre-trained vision-language models (VLMs), is a formidable challenge. Current Class Activation Map (CAM) methods highlight regions revealing the model's decision-making basis but lack clear saliency maps and detailed interpretability. To bridge this gap, we propose DecomCAM, a novel decomposition-and-integration method that distills shared patterns from channel activation maps. Utilizing singular value decomposition, DecomCAM decomposes class-discriminative activation maps into orthogonal sub-saliency maps (OSSMs), which are then integrated together based on their contribution to the target concept. Extensive experiments on six benchmarks reveal that DecomCAM not only excels in locating accuracy but also achieves an optimizing balance between interpretability and computational efficiency. Further analysis unveils that OSSMs correlate with discernible object components, facilitating a granular understanding of the model's reasoning. This positions DecomCAM as a potential tool for fine-grained interpretation of advanced deep learning models. The code is avaible at https://github.com/CapricornGuang/DecomCAM. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2405_18882 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | DecomCAM: Advancing Beyond Saliency Maps through Decomposition and Integration Yang, Yuguang Guo, Runtang Wu, Sheng Wang, Yimi Yang, Linlin Fan, Bo Zhong, Jilong Zhang, Juan Zhang, Baochang Computer Vision and Pattern Recognition Interpreting complex deep networks, notably pre-trained vision-language models (VLMs), is a formidable challenge. Current Class Activation Map (CAM) methods highlight regions revealing the model's decision-making basis but lack clear saliency maps and detailed interpretability. To bridge this gap, we propose DecomCAM, a novel decomposition-and-integration method that distills shared patterns from channel activation maps. Utilizing singular value decomposition, DecomCAM decomposes class-discriminative activation maps into orthogonal sub-saliency maps (OSSMs), which are then integrated together based on their contribution to the target concept. Extensive experiments on six benchmarks reveal that DecomCAM not only excels in locating accuracy but also achieves an optimizing balance between interpretability and computational efficiency. Further analysis unveils that OSSMs correlate with discernible object components, facilitating a granular understanding of the model's reasoning. This positions DecomCAM as a potential tool for fine-grained interpretation of advanced deep learning models. The code is avaible at https://github.com/CapricornGuang/DecomCAM. |
| title | DecomCAM: Advancing Beyond Saliency Maps through Decomposition and Integration |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2405.18882 |