Inverse-RLignment: Large Language Model Alignment from Demonstrations through Inverse Reinforcement Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sun, Hao, van der Schaar, Mihaela
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915121567105024
author Sun, Hao
van der Schaar, Mihaela
author_facet Sun, Hao
van der Schaar, Mihaela
contents Aligning Large Language Models (LLMs) is crucial for enhancing their safety and utility. However, existing methods, primarily based on preference datasets, face challenges such as noisy labels, high annotation costs, and privacy concerns. In this work, we introduce Alignment from Demonstrations (AfD), a novel approach leveraging high-quality demonstration data to overcome these challenges. We formalize AfD within a sequential decision-making framework, highlighting its unique challenge of missing reward signals. Drawing insights from forward and inverse reinforcement learning, we introduce divergence minimization objectives for AfD. Analytically, we elucidate the mass-covering and mode-seeking behaviors of various approaches, explaining when and why certain methods are superior. Practically, we propose a computationally efficient algorithm that extrapolates over a tailored reward model for AfD. We validate our key insights through experiments on the Harmless and Helpful tasks, demonstrating their strong empirical performance while maintaining simplicity.
format Preprint
id arxiv_https___arxiv_org_abs_2405_15624
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Inverse-RLignment: Large Language Model Alignment from Demonstrations through Inverse Reinforcement Learning
Sun, Hao
van der Schaar, Mihaela
Machine Learning
Artificial Intelligence
Aligning Large Language Models (LLMs) is crucial for enhancing their safety and utility. However, existing methods, primarily based on preference datasets, face challenges such as noisy labels, high annotation costs, and privacy concerns. In this work, we introduce Alignment from Demonstrations (AfD), a novel approach leveraging high-quality demonstration data to overcome these challenges. We formalize AfD within a sequential decision-making framework, highlighting its unique challenge of missing reward signals. Drawing insights from forward and inverse reinforcement learning, we introduce divergence minimization objectives for AfD. Analytically, we elucidate the mass-covering and mode-seeking behaviors of various approaches, explaining when and why certain methods are superior. Practically, we propose a computationally efficient algorithm that extrapolates over a tailored reward model for AfD. We validate our key insights through experiments on the Harmless and Helpful tasks, demonstrating their strong empirical performance while maintaining simplicity.
title Inverse-RLignment: Large Language Model Alignment from Demonstrations through Inverse Reinforcement Learning
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2405.15624