Learning Utilities from Demonstrations in Markov Decision Processes
Fuente:
arXiv
Saved in:
| Main Authors: | Lazzati, Filippo, Metelli, Alberto Maria |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Robustness in the Face of Partial Identifiability in Reward Learning
by: Lazzati, Filippo, et al.
Published: (2025)
by: Lazzati, Filippo, et al.
Published: (2025)
Imitation Learning as Return Distribution Matching
by: Lazzati, Filippo, et al.
Published: (2025)
by: Lazzati, Filippo, et al.
Published: (2025)
Generalizing Behavior via Inverse Reinforcement Learning with Closed-Form Reward Centroids
by: Lazzati, Filippo, et al.
Published: (2025)
by: Lazzati, Filippo, et al.
Published: (2025)
Reward Compatibility: A Framework for Inverse RL
by: Lazzati, Filippo, et al.
Published: (2025)
by: Lazzati, Filippo, et al.
Published: (2025)
Performance Improvement Bounds for Lipschitz Configurable Markov Decision Processes
by: Metelli, Alberto Maria
Published: (2024)
by: Metelli, Alberto Maria
Published: (2024)
Offline Inverse RL: New Solution Concepts and Provably Efficient Algorithms
by: Lazzati, Filippo, et al.
Published: (2024)
by: Lazzati, Filippo, et al.
Published: (2024)
How does Inverse RL Scale to Large State Spaces? A Provably Efficient Approach
by: Lazzati, Filippo, et al.
Published: (2024)
by: Lazzati, Filippo, et al.
Published: (2024)
Learning Deterministic Policies with Policy Gradients in Constrained Markov Decision Processes
by: Montenegro, Alessandro, et al.
Published: (2025)
by: Montenegro, Alessandro, et al.
Published: (2025)
The Number of Trials Matters in Infinite-Horizon General-Utility Markov Decision Processes
by: Santos, Pedro P., et al.
Published: (2024)
by: Santos, Pedro P., et al.
Published: (2024)
Solving General-Utility Markov Decision Processes in the Single-Trial Regime with Online Planning
by: Santos, Pedro P., et al.
Published: (2025)
by: Santos, Pedro P., et al.
Published: (2025)
Risk-sensitive Markov Decision Process and Learning under General Utility Functions
by: Wu, Zhengqi, et al.
Published: (2023)
by: Wu, Zhengqi, et al.
Published: (2023)
Learning Constrained Markov Decision Processes With Non-stationary Rewards and Constraints
by: Stradi, Francesco Emanuele, et al.
Published: (2024)
by: Stradi, Francesco Emanuele, et al.
Published: (2024)
Efficient Learning of POMDPs with Known Observation Model in Average-Reward Setting
by: Russo, Alessio, et al.
Published: (2024)
by: Russo, Alessio, et al.
Published: (2024)
Learning in Markov Decision Processes with Exogenous Dynamics
by: Maran, Davide, et al.
Published: (2026)
by: Maran, Davide, et al.
Published: (2026)
Interpetable Target-Feature Aggregation for Multi-Task Learning based on Bias-Variance Analysis
by: Bonetti, Paolo, et al.
Published: (2024)
by: Bonetti, Paolo, et al.
Published: (2024)
A Provably Efficient Option-Based Algorithm for both High-Level and Low-Level Learning
by: Drappo, Gianluca, et al.
Published: (2024)
by: Drappo, Gianluca, et al.
Published: (2024)
Monitored Markov Decision Processes
by: Parisi, Simone, et al.
Published: (2024)
by: Parisi, Simone, et al.
Published: (2024)
Federated Control in Markov Decision Processes
by: Jin, Hao, et al.
Published: (2024)
by: Jin, Hao, et al.
Published: (2024)
Learning Optimal Deterministic Policies with Stochastic Policy Gradients
by: Montenegro, Alessandro, et al.
Published: (2024)
by: Montenegro, Alessandro, et al.
Published: (2024)
Generalized Linear Markov Decision Process
by: Zhang, Sinian, et al.
Published: (2025)
by: Zhang, Sinian, et al.
Published: (2025)
Transition Transfer $Q$-Learning for Composite Markov Decision Processes
by: Chai, Jinhang, et al.
Published: (2025)
by: Chai, Jinhang, et al.
Published: (2025)
Learning Markov Decision Processes under Fully Bandit Feedback
by: Zhuo, Zhengjia, et al.
Published: (2026)
by: Zhuo, Zhengjia, et al.
Published: (2026)
Local Linearity: the Key for No-regret Reinforcement Learning in Continuous MDPs
by: Maran, Davide, et al.
Published: (2024)
by: Maran, Davide, et al.
Published: (2024)
Last-Iterate Global Convergence of Policy Gradients for Constrained Reinforcement Learning
by: Montenegro, Alessandro, et al.
Published: (2024)
by: Montenegro, Alessandro, et al.
Published: (2024)
Open Problem: Tight Bounds for Kernelized Multi-Armed Bandits with Bernoulli Rewards
by: Mussi, Marco, et al.
Published: (2024)
by: Mussi, Marco, et al.
Published: (2024)
Sliding-Window Thompson Sampling for Non-Stationary Settings
by: Fiandri, Marco, et al.
Published: (2024)
by: Fiandri, Marco, et al.
Published: (2024)
Achieving $\widetilde{\mathcal{O}}(\sqrt{T})$ Regret in Average-Reward POMDPs with Known Observation Models
by: Russo, Alessio, et al.
Published: (2025)
by: Russo, Alessio, et al.
Published: (2025)
Generalized Kernelized Bandits: A Novel Self-Normalized Bernstein-Like Dimension-Free Inequality and Regret Bounds
by: Metelli, Alberto Maria, et al.
Published: (2025)
by: Metelli, Alberto Maria, et al.
Published: (2025)
Pure Exploration under Mediators' Feedback
by: Poiani, Riccardo, et al.
Published: (2023)
by: Poiani, Riccardo, et al.
Published: (2023)
Thompson Sampling-like Algorithms for Stochastic Rising Bandits
by: Fiandri, Marco, et al.
Published: (2025)
by: Fiandri, Marco, et al.
Published: (2025)
A Refined Analysis of UCBVI
by: Drago, Simone, et al.
Published: (2025)
by: Drago, Simone, et al.
Published: (2025)
Interaction-Grounded Learning for Contextual Markov Decision Processes with Personalized Feedback
by: Zhang, Mengxiao, et al.
Published: (2026)
by: Zhang, Mengxiao, et al.
Published: (2026)
Impact of Markov Decision Process Design on Sim-to-Real Reinforcement Learning
by: Krau, Tatjana, et al.
Published: (2026)
by: Krau, Tatjana, et al.
Published: (2026)
No-Regret Reinforcement Learning in Smooth MDPs
by: Maran, Davide, et al.
Published: (2024)
by: Maran, Davide, et al.
Published: (2024)
Optimal Decision Tree Policies for Markov Decision Processes
by: Vos, Daniël, et al.
Published: (2023)
by: Vos, Daniël, et al.
Published: (2023)
Inverse Reinforcement Learning with Sub-optimal Experts
by: Poiani, Riccardo, et al.
Published: (2024)
by: Poiani, Riccardo, et al.
Published: (2024)
Policy Testing in Markov Decision Processes
by: Ariu, Kaito, et al.
Published: (2025)
by: Ariu, Kaito, et al.
Published: (2025)
Rising Rested Bandits: Lower Bounds and Efficient Algorithms
by: Fiandri, Marco, et al.
Published: (2024)
by: Fiandri, Marco, et al.
Published: (2024)
Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes
by: Wang, Zijian, et al.
Published: (2025)
by: Wang, Zijian, et al.
Published: (2025)
Dynamic Deep-Reinforcement-Learning Algorithm in Partially Observable Markov Decision Processes
by: Omi, Saki, et al.
Published: (2023)
by: Omi, Saki, et al.
Published: (2023)
Similar Items
-
Robustness in the Face of Partial Identifiability in Reward Learning
by: Lazzati, Filippo, et al.
Published: (2025) -
Imitation Learning as Return Distribution Matching
by: Lazzati, Filippo, et al.
Published: (2025) -
Generalizing Behavior via Inverse Reinforcement Learning with Closed-Form Reward Centroids
by: Lazzati, Filippo, et al.
Published: (2025) -
Reward Compatibility: A Framework for Inverse RL
by: Lazzati, Filippo, et al.
Published: (2025) -
Performance Improvement Bounds for Lipschitz Configurable Markov Decision Processes
by: Metelli, Alberto Maria
Published: (2024)