Bayesian Robust Optimization for Imitation Learning
Fuente:
arXiv
Guardado en:
| Autores principales: | Brown, Daniel S., Niekum, Scott, Petrik, Marek |
|---|---|
| Formato: | Preprint |
| Publicado: |
2020
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
On the Benefits of Inducing Local Lipschitzness for Robust Generative Adversarial Imitation Learning
por: Memarian, Farzan, et al.
Publicado: (2021)
por: Memarian, Farzan, et al.
Publicado: (2021)
Dual RL: Unification and New Methods for Reinforcement and Imitation Learning
por: Sikchi, Harshit, et al.
Publicado: (2023)
por: Sikchi, Harshit, et al.
Publicado: (2023)
A Dual Approach to Imitation Learning from Observations with Offline Datasets
por: Sikchi, Harshit, et al.
Publicado: (2024)
por: Sikchi, Harshit, et al.
Publicado: (2024)
Bayesian Regret Minimization in Offline Bandits
por: Petrik, Marek, et al.
Publicado: (2023)
por: Petrik, Marek, et al.
Publicado: (2023)
An Optimal Discriminator Weighted Imitation Perspective for Reinforcement Learning
por: Xu, Haoran, et al.
Publicado: (2025)
por: Xu, Haoran, et al.
Publicado: (2025)
Training ML Models with Predictable Failures
por: Schwarzer, Will, et al.
Publicado: (2026)
por: Schwarzer, Will, et al.
Publicado: (2026)
Policy Gradient in Robust MDPs with Global Convergence Guarantee
por: Wang, Qiuhao, et al.
Publicado: (2022)
por: Wang, Qiuhao, et al.
Publicado: (2022)
Percentile Criterion Optimization in Offline Reinforcement Learning
por: Lobo, Elita A., et al.
Publicado: (2024)
por: Lobo, Elita A., et al.
Publicado: (2024)
Solving Multi-Model MDPs by Coordinate Ascent and Dynamic Programming
por: Su, Xihong, et al.
Publicado: (2024)
por: Su, Xihong, et al.
Publicado: (2024)
Efficient Algorithms for Robust Markov Decision Processes with $s$-Rectangular Ambiguity Sets
por: Ho, Chin Pang, et al.
Publicado: (2026)
por: Ho, Chin Pang, et al.
Publicado: (2026)
Reducing Blackwell and Average Optimality to Discounted MDPs via the Blackwell Discount Factor
por: Grand-Clément, Julien, et al.
Publicado: (2023)
por: Grand-Clément, Julien, et al.
Publicado: (2023)
Evaluation-Aware Reinforcement Learning
por: Deshmukh, Shripad Vilasrao, et al.
Publicado: (2025)
por: Deshmukh, Shripad Vilasrao, et al.
Publicado: (2025)
Policy Gradient for Robust Markov Decision Processes
por: Wang, Qiuhao, et al.
Publicado: (2024)
por: Wang, Qiuhao, et al.
Publicado: (2024)
RASR: Risk-Averse Soft-Robust MDPs with EVaR and Entropic Risk
por: Hau, Jia Lin, et al.
Publicado: (2022)
por: Hau, Jia Lin, et al.
Publicado: (2022)
Supervised Reward Inference
por: Schwarzer, Will, et al.
Publicado: (2025)
por: Schwarzer, Will, et al.
Publicado: (2025)
Risk-Averse Total-Reward Reinforcement Learning
por: Su, Xihong, et al.
Publicado: (2025)
por: Su, Xihong, et al.
Publicado: (2025)
Risk-averse Total-reward MDPs with ERM and EVaR
por: Su, Xihong, et al.
Publicado: (2024)
por: Su, Xihong, et al.
Publicado: (2024)
Pareto-Optimal Learning from Preferences with Hidden Context
por: Bahlous-Boldi, Ryan, et al.
Publicado: (2024)
por: Bahlous-Boldi, Ryan, et al.
Publicado: (2024)
Pairwise or Pointwise? Evaluating Feedback Protocols for Bias in LLM-Based Evaluation
por: Tripathi, Tuhina, et al.
Publicado: (2025)
por: Tripathi, Tuhina, et al.
Publicado: (2025)
Safe RLHF Beyond Expectation: Stochastic Dominance for Universal Spectral Risk Control
por: Chittepu, Yaswanth, et al.
Publicado: (2026)
por: Chittepu, Yaswanth, et al.
Publicado: (2026)
Robust Imitation Learning for Automated Game Testing
por: Amadori, Pierluigi Vito, et al.
Publicado: (2024)
por: Amadori, Pierluigi Vito, et al.
Publicado: (2024)
Learning Action-based Representations Using Invariance
por: Rudolph, Max, et al.
Publicado: (2024)
por: Rudolph, Max, et al.
Publicado: (2024)
Q-learning for Quantile MDPs: A Decomposition, Performance, and Convergence Analysis
por: Hau, Jia Lin, et al.
Publicado: (2024)
por: Hau, Jia Lin, et al.
Publicado: (2024)
Adaptive Margin RLHF via Preference over Preferences
por: Chittepu, Yaswanth, et al.
Publicado: (2025)
por: Chittepu, Yaswanth, et al.
Publicado: (2025)
A Bayesian Solution To The Imitation Gap
por: Vuorio, Risto, et al.
Publicado: (2024)
por: Vuorio, Risto, et al.
Publicado: (2024)
SMORE: Score Models for Offline Goal-Conditioned Reinforcement Learning
por: Sikchi, Harshit, et al.
Publicado: (2023)
por: Sikchi, Harshit, et al.
Publicado: (2023)
Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints
por: Chittepu, Yaswanth, et al.
Publicado: (2025)
por: Chittepu, Yaswanth, et al.
Publicado: (2025)
Robust Offline Imitation Learning from Diverse Auxiliary Data
por: Ghosh, Udita, et al.
Publicado: (2024)
por: Ghosh, Udita, et al.
Publicado: (2024)
Autonomous Assessment of Demonstration Sufficiency via Bayesian Inverse Reinforcement Learning
por: Trinh, Tu, et al.
Publicado: (2022)
por: Trinh, Tu, et al.
Publicado: (2022)
Contrastive Preference Learning: Learning from Human Feedback without RL
por: Hejna, Joey, et al.
Publicado: (2023)
por: Hejna, Joey, et al.
Publicado: (2023)
Probabilistic Safety Guarantee for Stochastic Control Systems Using Average Reward MDPs
por: Omidi, Saber, et al.
Publicado: (2025)
por: Omidi, Saber, et al.
Publicado: (2025)
Null Counterfactual Factor Interactions for Goal-Conditioned Reinforcement Learning
por: Chuck, Caleb, et al.
Publicado: (2025)
por: Chuck, Caleb, et al.
Publicado: (2025)
DeepForge: Leveraging AI for Microstructural Control in Metal Forming via Model Predictive Control
por: Petrik, Jan, et al.
Publicado: (2024)
por: Petrik, Jan, et al.
Publicado: (2024)
Data Poisoning Attacks on Off-Policy Policy Evaluation Methods
por: Lobo, Elita, et al.
Publicado: (2024)
por: Lobo, Elita, et al.
Publicado: (2024)
A Descriptive and Normative Theory of Human Beliefs in RLHF
por: Dandekar, Sylee, et al.
Publicado: (2025)
por: Dandekar, Sylee, et al.
Publicado: (2025)
Blending Imitation and Reinforcement Learning for Robust Policy Improvement
por: Liu, Xuefeng, et al.
Publicado: (2023)
por: Liu, Xuefeng, et al.
Publicado: (2023)
Deconfounding Imitation Learning with Variational Inference
por: Vuorio, Risto, et al.
Publicado: (2022)
por: Vuorio, Risto, et al.
Publicado: (2022)
Balance Equation-based Distributionally Robust Offline Imitation Learning
por: Agrawal, Rishabh, et al.
Publicado: (2025)
por: Agrawal, Rishabh, et al.
Publicado: (2025)
Imitation Learning from Observation through Optimal Transport
por: Chang, Wei-Di, et al.
Publicado: (2023)
por: Chang, Wei-Di, et al.
Publicado: (2023)
Robust Entropy Search for Safe Efficient Bayesian Optimization
por: Weichert, Dorina, et al.
Publicado: (2024)
por: Weichert, Dorina, et al.
Publicado: (2024)
Ejemplares similares
-
On the Benefits of Inducing Local Lipschitzness for Robust Generative Adversarial Imitation Learning
por: Memarian, Farzan, et al.
Publicado: (2021) -
Dual RL: Unification and New Methods for Reinforcement and Imitation Learning
por: Sikchi, Harshit, et al.
Publicado: (2023) -
A Dual Approach to Imitation Learning from Observations with Offline Datasets
por: Sikchi, Harshit, et al.
Publicado: (2024) -
Bayesian Regret Minimization in Offline Bandits
por: Petrik, Marek, et al.
Publicado: (2023) -
An Optimal Discriminator Weighted Imitation Perspective for Reinforcement Learning
por: Xu, Haoran, et al.
Publicado: (2025)