Beyond Patches: Mining Interpretable Part-Prototypes for Explainable AI

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Alehdaghi, Mahdi, Bhattacharya, Rajarshi, Shamsolmoali, Pourya, Cruz, Rafael M. O., Heritier, Maguelonne, Granger, Eric
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917094656835584
author Alehdaghi, Mahdi
Bhattacharya, Rajarshi
Shamsolmoali, Pourya
Cruz, Rafael M. O.
Heritier, Maguelonne
Granger, Eric
author_facet Alehdaghi, Mahdi
Bhattacharya, Rajarshi
Shamsolmoali, Pourya
Cruz, Rafael M. O.
Heritier, Maguelonne
Granger, Eric
contents As AI systems grow more capable, it becomes increasingly important that their decisions remain understandable and aligned with human expectations. A key challenge is the limited interpretability of deep models. Post-hoc methods like GradCAM offer heatmaps but provide limited conceptual insight, while prototype-based approaches offer example-based explanations but often rely on rigid region selection and lack semantic consistency. To address these limitations, we propose PCMNet, a part-prototypical concept mining network that learns human-comprehensible prototypes from meaningful image regions without additional supervision. By clustering these prototypes into concept groups and extracting concept activation vectors, PCMNet provides structured, concept-level explanations and enhances robustness to occlusion and challenging conditions, which are both critical for building reliable and aligned AI systems. Experiments across multiple image classification benchmarks show that PCMNet outperforms state-of-the-art methods in interpretability, stability, and robustness. This work contributes to AI alignment by enhancing transparency, controllability, and trustworthiness in AI systems. Our code is available at: https://github.com/alehdaghi/PCMNet.
format Preprint
id arxiv_https___arxiv_org_abs_2504_12197
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Beyond Patches: Mining Interpretable Part-Prototypes for Explainable AI
Alehdaghi, Mahdi
Bhattacharya, Rajarshi
Shamsolmoali, Pourya
Cruz, Rafael M. O.
Heritier, Maguelonne
Granger, Eric
Computer Vision and Pattern Recognition
As AI systems grow more capable, it becomes increasingly important that their decisions remain understandable and aligned with human expectations. A key challenge is the limited interpretability of deep models. Post-hoc methods like GradCAM offer heatmaps but provide limited conceptual insight, while prototype-based approaches offer example-based explanations but often rely on rigid region selection and lack semantic consistency. To address these limitations, we propose PCMNet, a part-prototypical concept mining network that learns human-comprehensible prototypes from meaningful image regions without additional supervision. By clustering these prototypes into concept groups and extracting concept activation vectors, PCMNet provides structured, concept-level explanations and enhances robustness to occlusion and challenging conditions, which are both critical for building reliable and aligned AI systems. Experiments across multiple image classification benchmarks show that PCMNet outperforms state-of-the-art methods in interpretability, stability, and robustness. This work contributes to AI alignment by enhancing transparency, controllability, and trustworthiness in AI systems. Our code is available at: https://github.com/alehdaghi/PCMNet.
title Beyond Patches: Mining Interpretable Part-Prototypes for Explainable AI
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2504.12197