PAGAR: Taming Reward Misalignment in Inverse Reinforcement Learning-Based Imitation Learning with Protagonist Antagonist Guided Adversarial Reward
Fuente:
arXiv
Saved in:
| Main Authors: | Zhou, Weichao, Li, Wenchao |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Rethinking Inverse Reinforcement Learning: from Data Alignment to Task Alignment
by: Zhou, Weichao, et al.
Published: (2024)
by: Zhou, Weichao, et al.
Published: (2024)
Rethinking Adversarial Inverse Reinforcement Learning: Policy Imitation, Transferable Reward Recovery and Algebraic Equilibrium Proof
by: Zhang, Yangchun, et al.
Published: (2024)
by: Zhang, Yangchun, et al.
Published: (2024)
Diffusion-Reward Adversarial Imitation Learning
by: Lai, Chun-Mao, et al.
Published: (2024)
by: Lai, Chun-Mao, et al.
Published: (2024)
On Reward Transferability in Adversarial Inverse Reinforcement Learning: Insights from Random Matrix Theory
by: Zhang, Yangchun, et al.
Published: (2024)
by: Zhang, Yangchun, et al.
Published: (2024)
Learning Rewards, Not Labels: Adversarial Inverse Reinforcement Learning for Machinery Fault Detection
by: Neupane, Dhiraj, et al.
Published: (2026)
by: Neupane, Dhiraj, et al.
Published: (2026)
Temporal Logic Specification-Conditioned Decision Transformer for Offline Safe Reinforcement Learning
by: Guo, Zijian, et al.
Published: (2024)
by: Guo, Zijian, et al.
Published: (2024)
On Feasible Rewards in Multi-Agent Inverse Reinforcement Learning
by: Freihaut, Till, et al.
Published: (2024)
by: Freihaut, Till, et al.
Published: (2024)
Bayesian Inverse Reinforcement Learning for Non-Markovian Rewards
by: Topper, Noah, et al.
Published: (2024)
by: Topper, Noah, et al.
Published: (2024)
From Novelty to Imitation: Self-Distilled Rewards for Offline Reinforcement Learning
by: Chaudhary, Gaurav, et al.
Published: (2025)
by: Chaudhary, Gaurav, et al.
Published: (2025)
Focal Reward: Balanced Reinforcement Learning under Rubric-Based Rewards
by: Huang, Yu, et al.
Published: (2026)
by: Huang, Yu, et al.
Published: (2026)
Reward-free World Models for Online Imitation Learning
by: Li, Shangzhe, et al.
Published: (2024)
by: Li, Shangzhe, et al.
Published: (2024)
Multi Task Inverse Reinforcement Learning for Common Sense Reward
by: Glazer, Neta, et al.
Published: (2024)
by: Glazer, Neta, et al.
Published: (2024)
Expert Proximity as Surrogate Rewards for Single Demonstration Imitation Learning
by: Chiang, Chia-Cheng, et al.
Published: (2024)
by: Chiang, Chia-Cheng, et al.
Published: (2024)
Binary Reward Labeling: Bridging Offline Preference and Reward-Based Reinforcement Learning
by: Xu, Yinglun, et al.
Published: (2024)
by: Xu, Yinglun, et al.
Published: (2024)
TW-CRL: Time-Weighted Contrastive Reward Learning for Efficient Inverse Reinforcement Learning
by: Li, Yuxuan, et al.
Published: (2025)
by: Li, Yuxuan, et al.
Published: (2025)
Parent-Guided Semantic Reward Model (PGSRM): Embedding-Based Reward Functions for Reinforcement Learning of Transformer Language Models
by: Plashchinsky, Alexandr
Published: (2025)
by: Plashchinsky, Alexandr
Published: (2025)
Off-Dynamics Reinforcement Learning via Domain Adaptation and Reward Augmented Imitation
by: Guo, Yihong, et al.
Published: (2024)
by: Guo, Yihong, et al.
Published: (2024)
Reward-Conditioned Reinforcement Learning
by: Nauman, Michal, et al.
Published: (2026)
by: Nauman, Michal, et al.
Published: (2026)
Reward Transfer from Inverse Reinforcement Learning: A Coupled Minimax Approach
by: Hao, Guang-Yuan, et al.
Published: (2026)
by: Hao, Guang-Yuan, et al.
Published: (2026)
Inverse Contextual Bandits without Rewards: Learning from a Non-Stationary Learner via Suffix Imitation
by: Kong, Yuqi, et al.
Published: (2026)
by: Kong, Yuqi, et al.
Published: (2026)
Inverse Reinforcement Learning with Dynamic Reward Scaling for LLM Alignment
by: Cheng, Ruoxi, et al.
Published: (2025)
by: Cheng, Ruoxi, et al.
Published: (2025)
Towards the Transferability of Rewards Recovered via Regularized Inverse Reinforcement Learning
by: Schlaginhaufen, Andreas, et al.
Published: (2024)
by: Schlaginhaufen, Andreas, et al.
Published: (2024)
Preference-Guided Learning for Sparse-Reward Multi-Agent Reinforcement Learning
by: Bui, The Viet, et al.
Published: (2025)
by: Bui, The Viet, et al.
Published: (2025)
The Distributional Reward Critic Framework for Reinforcement Learning Under Perturbed Rewards
by: Chen, Xi, et al.
Published: (2024)
by: Chen, Xi, et al.
Published: (2024)
Reinforcement Learning from Bagged Reward
by: Tang, Yuting, et al.
Published: (2024)
by: Tang, Yuting, et al.
Published: (2024)
The Value of Reward Lookahead in Reinforcement Learning
by: Merlis, Nadav, et al.
Published: (2024)
by: Merlis, Nadav, et al.
Published: (2024)
Informativeness of Reward Functions in Reinforcement Learning
by: Devidze, Rati, et al.
Published: (2024)
by: Devidze, Rati, et al.
Published: (2024)
Reward Design for Reinforcement Learning Agents
by: Devidze, Rati
Published: (2025)
by: Devidze, Rati
Published: (2025)
To the Max: Reinventing Reward in Reinforcement Learning
by: Veviurko, Grigorii, et al.
Published: (2024)
by: Veviurko, Grigorii, et al.
Published: (2024)
Generalizing Behavior via Inverse Reinforcement Learning with Closed-Form Reward Centroids
by: Lazzati, Filippo, et al.
Published: (2025)
by: Lazzati, Filippo, et al.
Published: (2025)
Inverse Reinforcement Learning with Switching Rewards and History Dependency for Characterizing Animal Behaviors
by: Ke, Jingyang, et al.
Published: (2025)
by: Ke, Jingyang, et al.
Published: (2025)
Enhancing Inverse Reinforcement Learning through Encoding Dynamic Information in Reward Shaping
by: Zhan, Simon Sinong, et al.
Published: (2024)
by: Zhan, Simon Sinong, et al.
Published: (2024)
Reward-Relevance-Filtered Linear Offline Reinforcement Learning
by: Zhou, Angela
Published: (2024)
by: Zhou, Angela
Published: (2024)
Reward-Zero: Language Embedding Driven Implicit Reward Mechanisms for Reinforcement Learning
by: Zhang, Heng, et al.
Published: (2026)
by: Zhang, Heng, et al.
Published: (2026)
Extrinsicaly Rewarded Soft Q Imitation Learning with Discriminator
by: Furuyama, Ryoma, et al.
Published: (2024)
by: Furuyama, Ryoma, et al.
Published: (2024)
Improving the Effectiveness of Potential-Based Reward Shaping in Reinforcement Learning
by: Müller, Henrik, et al.
Published: (2025)
by: Müller, Henrik, et al.
Published: (2025)
Constraints as Rewards: Reinforcement Learning for Robots without Reward Functions
by: Ishihara, Yu, et al.
Published: (2025)
by: Ishihara, Yu, et al.
Published: (2025)
Multi-Objective and Mixed-Reward Reinforcement Learning via Reward-Decorrelated Policy Optimization
by: Bai, Yang, et al.
Published: (2026)
by: Bai, Yang, et al.
Published: (2026)
Deep Reinforcement Learning with Hybrid Intrinsic Reward Model
by: Yuan, Mingqi, et al.
Published: (2025)
by: Yuan, Mingqi, et al.
Published: (2025)
Shrinking the Variance: Shrinkage Baselines for Reinforcement Learning with Verifiable Rewards
by: Zeng, Guanning, et al.
Published: (2025)
by: Zeng, Guanning, et al.
Published: (2025)
Similar Items
-
Rethinking Inverse Reinforcement Learning: from Data Alignment to Task Alignment
by: Zhou, Weichao, et al.
Published: (2024) -
Rethinking Adversarial Inverse Reinforcement Learning: Policy Imitation, Transferable Reward Recovery and Algebraic Equilibrium Proof
by: Zhang, Yangchun, et al.
Published: (2024) -
Diffusion-Reward Adversarial Imitation Learning
by: Lai, Chun-Mao, et al.
Published: (2024) -
On Reward Transferability in Adversarial Inverse Reinforcement Learning: Insights from Random Matrix Theory
by: Zhang, Yangchun, et al.
Published: (2024) -
Learning Rewards, Not Labels: Adversarial Inverse Reinforcement Learning for Machinery Fault Detection
by: Neupane, Dhiraj, et al.
Published: (2026)