Saved in:
| Main Authors: | , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2602.14844 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917361386258432 |
|---|---|
| author | Malomgré, Elias Simoens, Pieter |
| author_facet | Malomgré, Elias Simoens, Pieter |
| contents | AI alignment is growing in importance, yet many current approaches learn safety behavior by directly modifying policy parameters, entangling normative constraints with the underlying policy. This often yields opaque, difficult-to-edit alignment artifacts and reduces their reuse across models or deployments, a failure mode we term Alignment Waste. We propose Interactionless Inverse Reinforcement Learning, a framework for learning inspectable, editable, and reusable reward artifacts separately from policy optimization. We further introduce the Alignment Flywheel, a human-in-the-loop lifecycle for iteratively auditing, patching, and hardening these artifacts through automated evaluation and refinement. Together, these ideas recast alignment from a disposable training expense into a durable, verifiable engineering asset. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2602_14844 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Interactionless Inverse Reinforcement Learning: A Data-Centric Framework for Durable Alignment Malomgré, Elias Simoens, Pieter Machine Learning AI alignment is growing in importance, yet many current approaches learn safety behavior by directly modifying policy parameters, entangling normative constraints with the underlying policy. This often yields opaque, difficult-to-edit alignment artifacts and reduces their reuse across models or deployments, a failure mode we term Alignment Waste. We propose Interactionless Inverse Reinforcement Learning, a framework for learning inspectable, editable, and reusable reward artifacts separately from policy optimization. We further introduce the Alignment Flywheel, a human-in-the-loop lifecycle for iteratively auditing, patching, and hardening these artifacts through automated evaluation and refinement. Together, these ideas recast alignment from a disposable training expense into a durable, verifiable engineering asset. |
| title | Interactionless Inverse Reinforcement Learning: A Data-Centric Framework for Durable Alignment |
| topic | Machine Learning |
| url | https://arxiv.org/abs/2602.14844 |