Interpretable Concept Bottlenecks to Align Reinforcement Learning Agents

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Delfosse, Quentin, Sztwiertnia, Sebastian, Rothermel, Mark, Stammer, Wolfgang, Kersting, Kristian
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910673672339456
author Delfosse, Quentin
Sztwiertnia, Sebastian
Rothermel, Mark
Stammer, Wolfgang
Kersting, Kristian
author_facet Delfosse, Quentin
Sztwiertnia, Sebastian
Rothermel, Mark
Stammer, Wolfgang
Kersting, Kristian
contents Goal misalignment, reward sparsity and difficult credit assignment are only a few of the many issues that make it difficult for deep reinforcement learning (RL) agents to learn optimal policies. Unfortunately, the black-box nature of deep neural networks impedes the inclusion of domain experts for inspecting the model and revising suboptimal policies. To this end, we introduce *Successive Concept Bottleneck Agents* (SCoBots), that integrate consecutive concept bottleneck (CB) layers. In contrast to current CB models, SCoBots do not just represent concepts as properties of individual objects, but also as relations between objects which is crucial for many RL tasks. Our experimental results provide evidence of SCoBots' competitive performances, but also of their potential for domain experts to understand and regularize their behavior. Among other things, SCoBots enabled us to identify a previously unknown misalignment problem in the iconic video game, Pong, and resolve it. Overall, SCoBots thus result in more human-aligned RL agents. Our code is available at https://github.com/k4ntz/SCoBots .
format Preprint
id arxiv_https___arxiv_org_abs_2401_05821
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Interpretable Concept Bottlenecks to Align Reinforcement Learning Agents
Delfosse, Quentin
Sztwiertnia, Sebastian
Rothermel, Mark
Stammer, Wolfgang
Kersting, Kristian
Machine Learning
Symbolic Computation
Goal misalignment, reward sparsity and difficult credit assignment are only a few of the many issues that make it difficult for deep reinforcement learning (RL) agents to learn optimal policies. Unfortunately, the black-box nature of deep neural networks impedes the inclusion of domain experts for inspecting the model and revising suboptimal policies. To this end, we introduce *Successive Concept Bottleneck Agents* (SCoBots), that integrate consecutive concept bottleneck (CB) layers. In contrast to current CB models, SCoBots do not just represent concepts as properties of individual objects, but also as relations between objects which is crucial for many RL tasks. Our experimental results provide evidence of SCoBots' competitive performances, but also of their potential for domain experts to understand and regularize their behavior. Among other things, SCoBots enabled us to identify a previously unknown misalignment problem in the iconic video game, Pong, and resolve it. Overall, SCoBots thus result in more human-aligned RL agents. Our code is available at https://github.com/k4ntz/SCoBots .
title Interpretable Concept Bottlenecks to Align Reinforcement Learning Agents
topic Machine Learning
Symbolic Computation
url https://arxiv.org/abs/2401.05821