Learning by Self-Explaining

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Stammer, Wolfgang, Friedrich, Felix, Steinmann, David, Brack, Manuel, Shindo, Hikaru, Kersting, Kristian
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914950863126528
author Stammer, Wolfgang
Friedrich, Felix
Steinmann, David
Brack, Manuel
Shindo, Hikaru
Kersting, Kristian
author_facet Stammer, Wolfgang
Friedrich, Felix
Steinmann, David
Brack, Manuel
Shindo, Hikaru
Kersting, Kristian
contents Much of explainable AI research treats explanations as a means for model inspection. Yet, this neglects findings from human psychology that describe the benefit of self-explanations in an agent's learning process. Motivated by this, we introduce a novel workflow in the context of image classification, termed Learning by Self-Explaining (LSX). LSX utilizes aspects of self-refining AI and human-guided explanatory machine learning. The underlying idea is that a learner model, in addition to optimizing for the original predictive task, is further optimized based on explanatory feedback from an internal critic model. Intuitively, a learner's explanations are considered "useful" if the internal critic can perform the same task given these explanations. We provide an overview of important components of LSX and, based on this, perform extensive experimental evaluations via three different example instantiations. Our results indicate improvements via Learning by Self-Explaining on several levels: in terms of model generalization, reducing the influence of confounding factors, and providing more task-relevant and faithful model explanations. Overall, our work provides evidence for the potential of self-explaining within the learning phase of an AI model.
format Preprint
id arxiv_https___arxiv_org_abs_2309_08395
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Learning by Self-Explaining
Stammer, Wolfgang
Friedrich, Felix
Steinmann, David
Brack, Manuel
Shindo, Hikaru
Kersting, Kristian
Artificial Intelligence
Machine Learning
Much of explainable AI research treats explanations as a means for model inspection. Yet, this neglects findings from human psychology that describe the benefit of self-explanations in an agent's learning process. Motivated by this, we introduce a novel workflow in the context of image classification, termed Learning by Self-Explaining (LSX). LSX utilizes aspects of self-refining AI and human-guided explanatory machine learning. The underlying idea is that a learner model, in addition to optimizing for the original predictive task, is further optimized based on explanatory feedback from an internal critic model. Intuitively, a learner's explanations are considered "useful" if the internal critic can perform the same task given these explanations. We provide an overview of important components of LSX and, based on this, perform extensive experimental evaluations via three different example instantiations. Our results indicate improvements via Learning by Self-Explaining on several levels: in terms of model generalization, reducing the influence of confounding factors, and providing more task-relevant and faithful model explanations. Overall, our work provides evidence for the potential of self-explaining within the learning phase of an AI model.
title Learning by Self-Explaining
topic Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2309.08395