Prototype-Grounded Concept Models for Verifiable Concept Alignment

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Colamonaco, Stefano, Debot, David, Barbiero, Pietro, Marra, Giuseppe
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911704035622912
author Colamonaco, Stefano
Debot, David
Barbiero, Pietro
Marra, Giuseppe
author_facet Colamonaco, Stefano
Debot, David
Barbiero, Pietro
Marra, Giuseppe
contents Concept Bottleneck Models (CBMs) aim to improve interpretability in Deep Learning by structuring predictions through human-understandable concepts, but they provide no way to verify whether learned concepts align with the human's intended meaning, hurting interpretability. We introduce Prototype-Grounded Concept Models (PGCMs), which ground concepts in learned visual prototypes: image parts that serve as explicit evidence for the concepts. This grounding enables direct inspection of concept semantics and supports targeted human intervention at the prototype level to correct misalignments. Empirically, PGCMs achieve similar predictive performance as state-of-the-art CBMs while substantially improving transparency, interpretability, and intervenability.
format Preprint
id arxiv_https___arxiv_org_abs_2604_16076
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Prototype-Grounded Concept Models for Verifiable Concept Alignment
Colamonaco, Stefano
Debot, David
Barbiero, Pietro
Marra, Giuseppe
Machine Learning
Artificial Intelligence
Neural and Evolutionary Computing
Concept Bottleneck Models (CBMs) aim to improve interpretability in Deep Learning by structuring predictions through human-understandable concepts, but they provide no way to verify whether learned concepts align with the human's intended meaning, hurting interpretability. We introduce Prototype-Grounded Concept Models (PGCMs), which ground concepts in learned visual prototypes: image parts that serve as explicit evidence for the concepts. This grounding enables direct inspection of concept semantics and supports targeted human intervention at the prototype level to correct misalignments. Empirically, PGCMs achieve similar predictive performance as state-of-the-art CBMs while substantially improving transparency, interpretability, and intervenability.
title Prototype-Grounded Concept Models for Verifiable Concept Alignment
topic Machine Learning
Artificial Intelligence
Neural and Evolutionary Computing
url https://arxiv.org/abs/2604.16076