Saved in:
Bibliographic Details
Main Authors: Malomgré, Elias, Simoens, Pieter
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2602.14844
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917361386258432
author Malomgré, Elias
Simoens, Pieter
author_facet Malomgré, Elias
Simoens, Pieter
contents AI alignment is growing in importance, yet many current approaches learn safety behavior by directly modifying policy parameters, entangling normative constraints with the underlying policy. This often yields opaque, difficult-to-edit alignment artifacts and reduces their reuse across models or deployments, a failure mode we term Alignment Waste. We propose Interactionless Inverse Reinforcement Learning, a framework for learning inspectable, editable, and reusable reward artifacts separately from policy optimization. We further introduce the Alignment Flywheel, a human-in-the-loop lifecycle for iteratively auditing, patching, and hardening these artifacts through automated evaluation and refinement. Together, these ideas recast alignment from a disposable training expense into a durable, verifiable engineering asset.
format Preprint
id arxiv_https___arxiv_org_abs_2602_14844
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Interactionless Inverse Reinforcement Learning: A Data-Centric Framework for Durable Alignment
Malomgré, Elias
Simoens, Pieter
Machine Learning
AI alignment is growing in importance, yet many current approaches learn safety behavior by directly modifying policy parameters, entangling normative constraints with the underlying policy. This often yields opaque, difficult-to-edit alignment artifacts and reduces their reuse across models or deployments, a failure mode we term Alignment Waste. We propose Interactionless Inverse Reinforcement Learning, a framework for learning inspectable, editable, and reusable reward artifacts separately from policy optimization. We further introduce the Alignment Flywheel, a human-in-the-loop lifecycle for iteratively auditing, patching, and hardening these artifacts through automated evaluation and refinement. Together, these ideas recast alignment from a disposable training expense into a durable, verifiable engineering asset.
title Interactionless Inverse Reinforcement Learning: A Data-Centric Framework for Durable Alignment
topic Machine Learning
url https://arxiv.org/abs/2602.14844