Saved in:
| Main Authors: | Brown, Daniel S., Niekum, Scott, Petrik, Marek |
|---|---|
| Format: | Preprint |
| Published: |
2020
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2007.12315 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
On the Benefits of Inducing Local Lipschitzness for Robust Generative Adversarial Imitation Learning
by: Memarian, Farzan, et al.
Published: (2021)
by: Memarian, Farzan, et al.
Published: (2021)
Dual RL: Unification and New Methods for Reinforcement and Imitation Learning
by: Sikchi, Harshit, et al.
Published: (2023)
by: Sikchi, Harshit, et al.
Published: (2023)
A Dual Approach to Imitation Learning from Observations with Offline Datasets
by: Sikchi, Harshit, et al.
Published: (2024)
by: Sikchi, Harshit, et al.
Published: (2024)
An Optimal Discriminator Weighted Imitation Perspective for Reinforcement Learning
by: Xu, Haoran, et al.
Published: (2025)
by: Xu, Haoran, et al.
Published: (2025)
Bayesian Regret Minimization in Offline Bandits
by: Petrik, Marek, et al.
Published: (2023)
by: Petrik, Marek, et al.
Published: (2023)
Training ML Models with Predictable Failures
by: Schwarzer, Will, et al.
Published: (2026)
by: Schwarzer, Will, et al.
Published: (2026)
Policy Gradient in Robust MDPs with Global Convergence Guarantee
by: Wang, Qiuhao, et al.
Published: (2022)
by: Wang, Qiuhao, et al.
Published: (2022)
Percentile Criterion Optimization in Offline Reinforcement Learning
by: Lobo, Elita A., et al.
Published: (2024)
by: Lobo, Elita A., et al.
Published: (2024)
Solving Multi-Model MDPs by Coordinate Ascent and Dynamic Programming
by: Su, Xihong, et al.
Published: (2024)
by: Su, Xihong, et al.
Published: (2024)
Efficient Algorithms for Robust Markov Decision Processes with $s$-Rectangular Ambiguity Sets
by: Ho, Chin Pang, et al.
Published: (2026)
by: Ho, Chin Pang, et al.
Published: (2026)
Reducing Blackwell and Average Optimality to Discounted MDPs via the Blackwell Discount Factor
by: Grand-Clément, Julien, et al.
Published: (2023)
by: Grand-Clément, Julien, et al.
Published: (2023)
Evaluation-Aware Reinforcement Learning
by: Deshmukh, Shripad Vilasrao, et al.
Published: (2025)
by: Deshmukh, Shripad Vilasrao, et al.
Published: (2025)
Policy Gradient for Robust Markov Decision Processes
by: Wang, Qiuhao, et al.
Published: (2024)
by: Wang, Qiuhao, et al.
Published: (2024)
RASR: Risk-Averse Soft-Robust MDPs with EVaR and Entropic Risk
by: Hau, Jia Lin, et al.
Published: (2022)
by: Hau, Jia Lin, et al.
Published: (2022)
Supervised Reward Inference
by: Schwarzer, Will, et al.
Published: (2025)
by: Schwarzer, Will, et al.
Published: (2025)
Risk-averse Total-reward MDPs with ERM and EVaR
by: Su, Xihong, et al.
Published: (2024)
by: Su, Xihong, et al.
Published: (2024)
Pareto-Optimal Learning from Preferences with Hidden Context
by: Bahlous-Boldi, Ryan, et al.
Published: (2024)
by: Bahlous-Boldi, Ryan, et al.
Published: (2024)
Risk-Averse Total-Reward Reinforcement Learning
by: Su, Xihong, et al.
Published: (2025)
by: Su, Xihong, et al.
Published: (2025)
Pairwise or Pointwise? Evaluating Feedback Protocols for Bias in LLM-Based Evaluation
by: Tripathi, Tuhina, et al.
Published: (2025)
by: Tripathi, Tuhina, et al.
Published: (2025)
Safe RLHF Beyond Expectation: Stochastic Dominance for Universal Spectral Risk Control
by: Chittepu, Yaswanth, et al.
Published: (2026)
by: Chittepu, Yaswanth, et al.
Published: (2026)
Learning Action-based Representations Using Invariance
by: Rudolph, Max, et al.
Published: (2024)
by: Rudolph, Max, et al.
Published: (2024)
Adaptive Margin RLHF via Preference over Preferences
by: Chittepu, Yaswanth, et al.
Published: (2025)
by: Chittepu, Yaswanth, et al.
Published: (2025)
Q-learning for Quantile MDPs: A Decomposition, Performance, and Convergence Analysis
by: Hau, Jia Lin, et al.
Published: (2024)
by: Hau, Jia Lin, et al.
Published: (2024)
SMORE: Score Models for Offline Goal-Conditioned Reinforcement Learning
by: Sikchi, Harshit, et al.
Published: (2023)
by: Sikchi, Harshit, et al.
Published: (2023)
Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints
by: Chittepu, Yaswanth, et al.
Published: (2025)
by: Chittepu, Yaswanth, et al.
Published: (2025)
Probabilistic Safety Guarantee for Stochastic Control Systems Using Average Reward MDPs
by: Omidi, Saber, et al.
Published: (2025)
by: Omidi, Saber, et al.
Published: (2025)
Contrastive Preference Learning: Learning from Human Feedback without RL
by: Hejna, Joey, et al.
Published: (2023)
by: Hejna, Joey, et al.
Published: (2023)
Data Poisoning Attacks on Off-Policy Policy Evaluation Methods
by: Lobo, Elita, et al.
Published: (2024)
by: Lobo, Elita, et al.
Published: (2024)
Null Counterfactual Factor Interactions for Goal-Conditioned Reinforcement Learning
by: Chuck, Caleb, et al.
Published: (2025)
by: Chuck, Caleb, et al.
Published: (2025)
Robust Imitation Learning for Automated Game Testing
by: Amadori, Pierluigi Vito, et al.
Published: (2024)
by: Amadori, Pierluigi Vito, et al.
Published: (2024)
A Descriptive and Normative Theory of Human Beliefs in RLHF
by: Dandekar, Sylee, et al.
Published: (2025)
by: Dandekar, Sylee, et al.
Published: (2025)
A Bayesian Solution To The Imitation Gap
by: Vuorio, Risto, et al.
Published: (2024)
by: Vuorio, Risto, et al.
Published: (2024)
DeepForge: Leveraging AI for Microstructural Control in Metal Forming via Model Predictive Control
by: Petrik, Jan, et al.
Published: (2024)
by: Petrik, Jan, et al.
Published: (2024)
Autonomous Assessment of Demonstration Sufficiency via Bayesian Inverse Reinforcement Learning
by: Trinh, Tu, et al.
Published: (2022)
by: Trinh, Tu, et al.
Published: (2022)
Automated Discovery of Functional Actual Causes in Complex Environments
by: Chuck, Caleb, et al.
Published: (2024)
by: Chuck, Caleb, et al.
Published: (2024)
Regularized Latent Dynamics Prediction is a Strong Baseline For Behavioral Foundation Models
by: Jajoo, Pranaya, et al.
Published: (2026)
by: Jajoo, Pranaya, et al.
Published: (2026)
Robust Offline Imitation Learning from Diverse Auxiliary Data
by: Ghosh, Udita, et al.
Published: (2024)
by: Ghosh, Udita, et al.
Published: (2024)
Blending Imitation and Reinforcement Learning for Robust Policy Improvement
by: Liu, Xuefeng, et al.
Published: (2023)
by: Liu, Xuefeng, et al.
Published: (2023)
Deconfounding Imitation Learning with Variational Inference
by: Vuorio, Risto, et al.
Published: (2022)
by: Vuorio, Risto, et al.
Published: (2022)
Imitation Learning from Observation through Optimal Transport
by: Chang, Wei-Di, et al.
Published: (2023)
by: Chang, Wei-Di, et al.
Published: (2023)
Similar Items
-
On the Benefits of Inducing Local Lipschitzness for Robust Generative Adversarial Imitation Learning
by: Memarian, Farzan, et al.
Published: (2021) -
Dual RL: Unification and New Methods for Reinforcement and Imitation Learning
by: Sikchi, Harshit, et al.
Published: (2023) -
A Dual Approach to Imitation Learning from Observations with Offline Datasets
by: Sikchi, Harshit, et al.
Published: (2024) -
An Optimal Discriminator Weighted Imitation Perspective for Reinforcement Learning
by: Xu, Haoran, et al.
Published: (2025) -
Bayesian Regret Minimization in Offline Bandits
by: Petrik, Marek, et al.
Published: (2023)