Concepts' Information Bottleneck Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Galliamov, Karim, Kazmi, Syed M Ahsan, Khan, Adil, Rivera, Adín Ramírez
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912907206328320
author Galliamov, Karim
Kazmi, Syed M Ahsan
Khan, Adil
Rivera, Adín Ramírez
author_facet Galliamov, Karim
Kazmi, Syed M Ahsan
Khan, Adil
Rivera, Adín Ramírez
contents Concept Bottleneck Models (CBMs) aim to deliver interpretable predictions by routing decisions through a human-understandable concept layer, yet they often suffer reduced accuracy and concept leakage that undermines faithfulness. We introduce an explicit Information Bottleneck regularizer on the concept layer that penalizes $I(X;C)$ while preserving task-relevant information in $I(C;Y)$, encouraging minimal-sufficient concept representations. We derive two practical variants (a variational objective and an entropy-based surrogate) and integrate them into standard CBM training without architectural changes or additional supervision. Evaluated across six CBM families and three benchmarks, the IB-regularized models consistently outperform their vanilla counterparts. Information-plane analyses further corroborate the intended behavior. These results indicate that enforcing a minimal-sufficient concept bottleneck improves both predictive performance and the reliability of concept-level interventions. The proposed regularizer offers a theoretic-grounded, architecture-agnostic path to more faithful and intervenable CBMs, resolving prior evaluation inconsistencies by aligning training protocols and demonstrating robust gains across model families and datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2602_14626
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Concepts' Information Bottleneck Models
Galliamov, Karim
Kazmi, Syed M Ahsan
Khan, Adil
Rivera, Adín Ramírez
Machine Learning
Concept Bottleneck Models (CBMs) aim to deliver interpretable predictions by routing decisions through a human-understandable concept layer, yet they often suffer reduced accuracy and concept leakage that undermines faithfulness. We introduce an explicit Information Bottleneck regularizer on the concept layer that penalizes $I(X;C)$ while preserving task-relevant information in $I(C;Y)$, encouraging minimal-sufficient concept representations. We derive two practical variants (a variational objective and an entropy-based surrogate) and integrate them into standard CBM training without architectural changes or additional supervision. Evaluated across six CBM families and three benchmarks, the IB-regularized models consistently outperform their vanilla counterparts. Information-plane analyses further corroborate the intended behavior. These results indicate that enforcing a minimal-sufficient concept bottleneck improves both predictive performance and the reliability of concept-level interventions. The proposed regularizer offers a theoretic-grounded, architecture-agnostic path to more faithful and intervenable CBMs, resolving prior evaluation inconsistencies by aligning training protocols and demonstrating robust gains across model families and datasets.
title Concepts' Information Bottleneck Models
topic Machine Learning
url https://arxiv.org/abs/2602.14626