Actionable Interpretability Must Be Defined in Terms of Symmetries

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Barbiero, Pietro, Zarlenga, Mateo Espinosa, Giannini, Francesco, Termine, Alberto, Bonchi, Filippo, Jamnik, Mateja, Marra, Giuseppe
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912857093832704
author Barbiero, Pietro
Zarlenga, Mateo Espinosa
Giannini, Francesco
Termine, Alberto
Bonchi, Filippo
Jamnik, Mateja
Marra, Giuseppe
author_facet Barbiero, Pietro
Zarlenga, Mateo Espinosa
Giannini, Francesco
Termine, Alberto
Bonchi, Filippo
Jamnik, Mateja
Marra, Giuseppe
contents This paper argues that interpretability research in Artificial Intelligence (AI) is fundamentally ill-posed as existing definitions of interpretability fail to describe how interpretability can be formally tested or designed for. We posit that actionable definitions of interpretability must be formulated in terms of *symmetries* that inform model design and lead to testable conditions. Under a probabilistic view, we hypothesise that four symmetries (inference equivariance, information invariance, concept-closure invariance, and structural invariance) suffice to (i) formalise interpretable models as a subclass of probabilistic models, (ii) yield a unified formulation of interpretable inference (e.g., alignment, interventions, and counterfactuals) as a form of Bayesian inversion, and (iii) provide a formal framework to verify compliance with safety standards and regulations.
format Preprint
id arxiv_https___arxiv_org_abs_2601_12913
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Actionable Interpretability Must Be Defined in Terms of Symmetries
Barbiero, Pietro
Zarlenga, Mateo Espinosa
Giannini, Francesco
Termine, Alberto
Bonchi, Filippo
Jamnik, Mateja
Marra, Giuseppe
Artificial Intelligence
Machine Learning
Neural and Evolutionary Computing
This paper argues that interpretability research in Artificial Intelligence (AI) is fundamentally ill-posed as existing definitions of interpretability fail to describe how interpretability can be formally tested or designed for. We posit that actionable definitions of interpretability must be formulated in terms of *symmetries* that inform model design and lead to testable conditions. Under a probabilistic view, we hypothesise that four symmetries (inference equivariance, information invariance, concept-closure invariance, and structural invariance) suffice to (i) formalise interpretable models as a subclass of probabilistic models, (ii) yield a unified formulation of interpretable inference (e.g., alignment, interventions, and counterfactuals) as a form of Bayesian inversion, and (iii) provide a formal framework to verify compliance with safety standards and regulations.
title Actionable Interpretability Must Be Defined in Terms of Symmetries
topic Artificial Intelligence
Machine Learning
Neural and Evolutionary Computing
url https://arxiv.org/abs/2601.12913