STAIR: Addressing Stage Misalignment through Temporal-Aligned Preference Reinforcement Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Luan, Yao, Mu, Ni, Yang, Yiqin, Xu, Bo, Jia, Qing-Shan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911182157250560
author Luan, Yao
Mu, Ni
Yang, Yiqin
Xu, Bo
Jia, Qing-Shan
author_facet Luan, Yao
Mu, Ni
Yang, Yiqin
Xu, Bo
Jia, Qing-Shan
contents Preference-based reinforcement learning (PbRL) bypasses complex reward engineering by learning rewards directly from human preferences, enabling better alignment with human intentions. However, its effectiveness in multi-stage tasks, where agents sequentially perform sub-tasks (e.g., navigation, grasping), is limited by stage misalignment: Comparing segments from mismatched stages, such as movement versus manipulation, results in uninformative feedback, thus hindering policy learning. In this paper, we validate the stage misalignment issue through theoretical analysis and empirical experiments. To address this issue, we propose STage-AlIgned Reward learning (STAIR), which first learns a stage approximation based on temporal distance, then prioritizes comparisons within the same stage. Temporal distance is learned via contrastive learning, which groups temporally close states into coherent stages, without predefined task knowledge, and adapts dynamically to policy changes. Extensive experiments demonstrate STAIR's superiority in multi-stage tasks and competitive performance in single-stage tasks. Furthermore, human studies show that stages approximated by STAIR are consistent with human cognition, confirming its effectiveness in mitigating stage misalignment.
format Preprint
id arxiv_https___arxiv_org_abs_2509_23802
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle STAIR: Addressing Stage Misalignment through Temporal-Aligned Preference Reinforcement Learning
Luan, Yao
Mu, Ni
Yang, Yiqin
Xu, Bo
Jia, Qing-Shan
Machine Learning
Preference-based reinforcement learning (PbRL) bypasses complex reward engineering by learning rewards directly from human preferences, enabling better alignment with human intentions. However, its effectiveness in multi-stage tasks, where agents sequentially perform sub-tasks (e.g., navigation, grasping), is limited by stage misalignment: Comparing segments from mismatched stages, such as movement versus manipulation, results in uninformative feedback, thus hindering policy learning. In this paper, we validate the stage misalignment issue through theoretical analysis and empirical experiments. To address this issue, we propose STage-AlIgned Reward learning (STAIR), which first learns a stage approximation based on temporal distance, then prioritizes comparisons within the same stage. Temporal distance is learned via contrastive learning, which groups temporally close states into coherent stages, without predefined task knowledge, and adapts dynamically to policy changes. Extensive experiments demonstrate STAIR's superiority in multi-stage tasks and competitive performance in single-stage tasks. Furthermore, human studies show that stages approximated by STAIR are consistent with human cognition, confirming its effectiveness in mitigating stage misalignment.
title STAIR: Addressing Stage Misalignment through Temporal-Aligned Preference Reinforcement Learning
topic Machine Learning
url https://arxiv.org/abs/2509.23802