An Optimal Discriminator Weighted Imitation Perspective for Reinforcement Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Xu, Haoran, Li, Shuozhe, Sikchi, Harshit, Niekum, Scott, Zhang, Amy |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Dual RL: Unification and New Methods for Reinforcement and Imitation Learning
by: Sikchi, Harshit, et al.
Published: (2023)
by: Sikchi, Harshit, et al.
Published: (2023)
A Dual Approach to Imitation Learning from Observations with Offline Datasets
by: Sikchi, Harshit, et al.
Published: (2024)
by: Sikchi, Harshit, et al.
Published: (2024)
SMORE: Score Models for Offline Goal-Conditioned Reinforcement Learning
by: Sikchi, Harshit, et al.
Published: (2023)
by: Sikchi, Harshit, et al.
Published: (2023)
Regularized Latent Dynamics Prediction is a Strong Baseline For Behavioral Foundation Models
by: Jajoo, Pranaya, et al.
Published: (2026)
by: Jajoo, Pranaya, et al.
Published: (2026)
Contrastive Preference Learning: Learning from Human Feedback without RL
by: Hejna, Joey, et al.
Published: (2023)
by: Hejna, Joey, et al.
Published: (2023)
Fast Adaptation with Behavioral Foundation Models
by: Sikchi, Harshit, et al.
Published: (2025)
by: Sikchi, Harshit, et al.
Published: (2025)
Proto Successor Measure: Representing the Behavior Space of an RL Agent
by: Agarwal, Siddhant, et al.
Published: (2024)
by: Agarwal, Siddhant, et al.
Published: (2024)
RLZero: Direct Policy Inference from Language Without In-Domain Supervision
by: Sikchi, Harshit, et al.
Published: (2024)
by: Sikchi, Harshit, et al.
Published: (2024)
Evaluation-Aware Reinforcement Learning
by: Deshmukh, Shripad Vilasrao, et al.
Published: (2025)
by: Deshmukh, Shripad Vilasrao, et al.
Published: (2025)
Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms
by: Rafailov, Rafael, et al.
Published: (2024)
by: Rafailov, Rafael, et al.
Published: (2024)
Pareto-Optimal Learning from Preferences with Hidden Context
by: Bahlous-Boldi, Ryan, et al.
Published: (2024)
by: Bahlous-Boldi, Ryan, et al.
Published: (2024)
Null Counterfactual Factor Interactions for Goal-Conditioned Reinforcement Learning
by: Chuck, Caleb, et al.
Published: (2025)
by: Chuck, Caleb, et al.
Published: (2025)
Learning Action-based Representations Using Invariance
by: Rudolph, Max, et al.
Published: (2024)
by: Rudolph, Max, et al.
Published: (2024)
CARE-RFT: Confidence-Anchored Reinforcement Finetuning for Reliable Reasoning in Large Language Models
by: Li, Shuozhe, et al.
Published: (2026)
by: Li, Shuozhe, et al.
Published: (2026)
Reinforcement Learning via Value Gradient Flow
by: Xu, Haoran, et al.
Published: (2026)
by: Xu, Haoran, et al.
Published: (2026)
Robot Air Hockey: A Manipulation Testbed for Robot Learning with Reinforcement Learning
by: Chuck, Caleb, et al.
Published: (2024)
by: Chuck, Caleb, et al.
Published: (2024)
Diffusion-DICE: In-Sample Diffusion Guidance for Offline Reinforcement Learning
by: Mao, Liyuan, et al.
Published: (2024)
by: Mao, Liyuan, et al.
Published: (2024)
Imitation Learning from Observation through Optimal Transport
by: Chang, Wei-Di, et al.
Published: (2023)
by: Chang, Wei-Di, et al.
Published: (2023)
Automated Discovery of Functional Actual Causes in Complex Environments
by: Chuck, Caleb, et al.
Published: (2024)
by: Chuck, Caleb, et al.
Published: (2024)
Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints
by: Chittepu, Yaswanth, et al.
Published: (2025)
by: Chittepu, Yaswanth, et al.
Published: (2025)
Imitation Bootstrapped Reinforcement Learning
by: Hu, Hengyuan, et al.
Published: (2023)
by: Hu, Hengyuan, et al.
Published: (2023)
Safe RLHF Beyond Expectation: Stochastic Dominance for Universal Spectral Risk Control
by: Chittepu, Yaswanth, et al.
Published: (2026)
by: Chittepu, Yaswanth, et al.
Published: (2026)
A Learnable Wavelet Transformer for Long-Short Equity Trading and Risk-Adjusted Return Optimization
by: Li, Shuozhe, et al.
Published: (2026)
by: Li, Shuozhe, et al.
Published: (2026)
RILe: Reinforced Imitation Learning
by: Albaba, Mert, et al.
Published: (2024)
by: Albaba, Mert, et al.
Published: (2024)
Learning Robust Reasoning through Guided Adversarial Self-Play
by: Li, Shuozhe, et al.
Published: (2026)
by: Li, Shuozhe, et al.
Published: (2026)
Bayesian Robust Optimization for Imitation Learning
by: Brown, Daniel S., et al.
Published: (2020)
by: Brown, Daniel S., et al.
Published: (2020)
Reinforcement Learning via Implicit Imitation Guidance
by: Dong, Perry, et al.
Published: (2025)
by: Dong, Perry, et al.
Published: (2025)
Imitating Cost-Constrained Behaviors in Reinforcement Learning
by: Shao, Qian, et al.
Published: (2024)
by: Shao, Qian, et al.
Published: (2024)
Extrinsicaly Rewarded Soft Q Imitation Learning with Discriminator
by: Furuyama, Ryoma, et al.
Published: (2024)
by: Furuyama, Ryoma, et al.
Published: (2024)
Imitating from auxiliary imperfect demonstrations via Adversarial Density Weighted Regression
by: Zhang, Ziqi, et al.
Published: (2024)
by: Zhang, Ziqi, et al.
Published: (2024)
Physics-informed Imitative Reinforcement Learning for Real-world Driving
by: Zhou, Hang, et al.
Published: (2024)
by: Zhou, Hang, et al.
Published: (2024)
Towards General-Purpose Model-Free Reinforcement Learning
by: Fujimoto, Scott, et al.
Published: (2025)
by: Fujimoto, Scott, et al.
Published: (2025)
Adaptive Margin RLHF via Preference over Preferences
by: Chittepu, Yaswanth, et al.
Published: (2025)
by: Chittepu, Yaswanth, et al.
Published: (2025)
Diverse Policies Recovering via Pointwise Mutual Information Weighted Imitation Learning
by: Yang, Hanlin, et al.
Published: (2024)
by: Yang, Hanlin, et al.
Published: (2024)
Blending Imitation and Reinforcement Learning for Robust Policy Improvement
by: Liu, Xuefeng, et al.
Published: (2023)
by: Liu, Xuefeng, et al.
Published: (2023)
A Descriptive and Normative Theory of Human Beliefs in RLHF
by: Dandekar, Sylee, et al.
Published: (2025)
by: Dandekar, Sylee, et al.
Published: (2025)
Boolean Satisfiability via Imitation Learning
by: Zhang, Zewei, et al.
Published: (2025)
by: Zhang, Zewei, et al.
Published: (2025)
Cross-Domain Imitation Learning via Optimal Transport
by: Fickinger, Arnaud, et al.
Published: (2021)
by: Fickinger, Arnaud, et al.
Published: (2021)
TDMPBC: Self-Imitative Reinforcement Learning for Humanoid Robot Control
by: Zhuang, Zifeng, et al.
Published: (2025)
by: Zhuang, Zifeng, et al.
Published: (2025)
Align Your Intents: Offline Imitation Learning via Optimal Transport
by: Bobrin, Maksim, et al.
Published: (2024)
by: Bobrin, Maksim, et al.
Published: (2024)
Similar Items
-
Dual RL: Unification and New Methods for Reinforcement and Imitation Learning
by: Sikchi, Harshit, et al.
Published: (2023) -
A Dual Approach to Imitation Learning from Observations with Offline Datasets
by: Sikchi, Harshit, et al.
Published: (2024) -
SMORE: Score Models for Offline Goal-Conditioned Reinforcement Learning
by: Sikchi, Harshit, et al.
Published: (2023) -
Regularized Latent Dynamics Prediction is a Strong Baseline For Behavioral Foundation Models
by: Jajoo, Pranaya, et al.
Published: (2026) -
Contrastive Preference Learning: Learning from Human Feedback without RL
by: Hejna, Joey, et al.
Published: (2023)