Towards Fine-Grained and Verifiable Concept Bottleneck Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fang, Yingying, Xu, Haijie, Wu, Shuang, Anish, Mariathasan, Yang, Guang
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917493931507712
author Fang, Yingying
Xu, Haijie
Wu, Shuang
Anish, Mariathasan
Yang, Guang
author_facet Fang, Yingying
Xu, Haijie
Wu, Shuang
Anish, Mariathasan
Yang, Guang
contents Concept Bottleneck Models (CBMs) offer interpretable alternatives to black-box predictors by introducing human-relatable concepts before the final output. However, existing CBMs struggle to verify whether predicted concepts correspond to the correct visual evidence, limiting their reliability. We propose a fine-grained CBM framework that grounds each concept in localized visual evidence, enabling direct inspection of where and how concepts are encoded. This design allows users to interpret predictions and verify that the model learns intended concepts rather than spurious correlations. Experiments on medical imaging benchmarks show that our learned concept space is information-complete and achieves predictive performance comparable to standard CBMs, while substantially improving transparency. Unlike post-hoc attribution methods, our framework validates both the presence and correctness of concept representations, bridging interpretability with verifiability. Our approach enhances the trustworthiness of CBMs and establishes a principled mechanism for human-model interaction at the concept level, paving the way toward more reliable and clinically actionable concept-based learning systems.
format Preprint
id arxiv_https___arxiv_org_abs_2605_14210
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Towards Fine-Grained and Verifiable Concept Bottleneck Models
Fang, Yingying
Xu, Haijie
Wu, Shuang
Anish, Mariathasan
Yang, Guang
Machine Learning
Artificial Intelligence
Concept Bottleneck Models (CBMs) offer interpretable alternatives to black-box predictors by introducing human-relatable concepts before the final output. However, existing CBMs struggle to verify whether predicted concepts correspond to the correct visual evidence, limiting their reliability. We propose a fine-grained CBM framework that grounds each concept in localized visual evidence, enabling direct inspection of where and how concepts are encoded. This design allows users to interpret predictions and verify that the model learns intended concepts rather than spurious correlations. Experiments on medical imaging benchmarks show that our learned concept space is information-complete and achieves predictive performance comparable to standard CBMs, while substantially improving transparency. Unlike post-hoc attribution methods, our framework validates both the presence and correctness of concept representations, bridging interpretability with verifiability. Our approach enhances the trustworthiness of CBMs and establishes a principled mechanism for human-model interaction at the concept level, paving the way toward more reliable and clinically actionable concept-based learning systems.
title Towards Fine-Grained and Verifiable Concept Bottleneck Models
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2605.14210