Offline Inverse RL: New Solution Concepts and Provably Efficient Algorithms
Fuente:
arXiv
Guardado en:
| Autores principales: | Lazzati, Filippo, Mutti, Mirco, Metelli, Alberto Maria |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
How does Inverse RL Scale to Large State Spaces? A Provably Efficient Approach
por: Lazzati, Filippo, et al.
Publicado: (2024)
por: Lazzati, Filippo, et al.
Publicado: (2024)
Reward Compatibility: A Framework for Inverse RL
por: Lazzati, Filippo, et al.
Publicado: (2025)
por: Lazzati, Filippo, et al.
Publicado: (2025)
Generalizing Behavior via Inverse Reinforcement Learning with Closed-Form Reward Centroids
por: Lazzati, Filippo, et al.
Publicado: (2025)
por: Lazzati, Filippo, et al.
Publicado: (2025)
Learning Utilities from Demonstrations in Markov Decision Processes
por: Lazzati, Filippo, et al.
Publicado: (2024)
por: Lazzati, Filippo, et al.
Publicado: (2024)
Robustness in the Face of Partial Identifiability in Reward Learning
por: Lazzati, Filippo, et al.
Publicado: (2025)
por: Lazzati, Filippo, et al.
Publicado: (2025)
Imitation Learning as Return Distribution Matching
por: Lazzati, Filippo, et al.
Publicado: (2025)
por: Lazzati, Filippo, et al.
Publicado: (2025)
A Provably Efficient Option-Based Algorithm for both High-Level and Low-Level Learning
por: Drappo, Gianluca, et al.
Publicado: (2024)
por: Drappo, Gianluca, et al.
Publicado: (2024)
Test-Time Regret Minimization in Meta Reinforcement Learning
por: Mutti, Mirco, et al.
Publicado: (2024)
por: Mutti, Mirco, et al.
Publicado: (2024)
Rising Rested Bandits: Lower Bounds and Efficient Algorithms
por: Fiandri, Marco, et al.
Publicado: (2024)
por: Fiandri, Marco, et al.
Publicado: (2024)
Performance Improvement Bounds for Lipschitz Configurable Markov Decision Processes
por: Metelli, Alberto Maria
Publicado: (2024)
por: Metelli, Alberto Maria
Publicado: (2024)
Thompson Sampling-like Algorithms for Stochastic Rising Bandits
por: Fiandri, Marco, et al.
Publicado: (2025)
por: Fiandri, Marco, et al.
Publicado: (2025)
Inverse Reinforcement Learning with Sub-optimal Experts
por: Poiani, Riccardo, et al.
Publicado: (2024)
por: Poiani, Riccardo, et al.
Publicado: (2024)
Towards Principled Unsupervised Multi-Agent Reinforcement Learning
por: Zamboni, Riccardo, et al.
Publicado: (2025)
por: Zamboni, Riccardo, et al.
Publicado: (2025)
Provably Efficient Representation Selection in Low-rank Markov Decision Processes: From Online to Offline RL
por: Zhang, Weitong, et al.
Publicado: (2021)
por: Zhang, Weitong, et al.
Publicado: (2021)
Efficient Learning of POMDPs with Known Observation Model in Average-Reward Setting
por: Russo, Alessio, et al.
Publicado: (2024)
por: Russo, Alessio, et al.
Publicado: (2024)
The Limits of Pure Exploration in POMDPs: When the Observation Entropy is Enough
por: Zamboni, Riccardo, et al.
Publicado: (2024)
por: Zamboni, Riccardo, et al.
Publicado: (2024)
A Classification View on Meta Learning Bandits
por: Mutti, Mirco, et al.
Publicado: (2025)
por: Mutti, Mirco, et al.
Publicado: (2025)
Algorithmic Guarantees for Distilling Supervised and Offline RL Datasets
por: Gupta, Aaryan, et al.
Publicado: (2025)
por: Gupta, Aaryan, et al.
Publicado: (2025)
Enhancing Diversity in Parallel Agents: A Maximum State Entropy Exploration Story
por: De Paola, Vincenzo, et al.
Publicado: (2025)
por: De Paola, Vincenzo, et al.
Publicado: (2025)
K-Myriad: Jump-starting reinforcement learning with unsupervised parallel agents
por: De Paola, Vincenzo, et al.
Publicado: (2026)
por: De Paola, Vincenzo, et al.
Publicado: (2026)
Provably Efficient Offline-to-Online Value Adaptation with General Function Approximation
por: Li, Shangzhe, et al.
Publicado: (2026)
por: Li, Shangzhe, et al.
Publicado: (2026)
A Theoretical Framework for Partially Observed Reward-States in RLHF
por: Kausik, Chinmaya, et al.
Publicado: (2024)
por: Kausik, Chinmaya, et al.
Publicado: (2024)
How to Explore with Belief: State Entropy Maximization in POMDPs
por: Zamboni, Riccardo, et al.
Publicado: (2024)
por: Zamboni, Riccardo, et al.
Publicado: (2024)
From Parameters to Behaviors: Unsupervised Compression of the Policy Space
por: Tenedini, Davide, et al.
Publicado: (2025)
por: Tenedini, Davide, et al.
Publicado: (2025)
Query-Dependent Prompt Evaluation and Optimization with Offline Inverse RL
por: Sun, Hao, et al.
Publicado: (2023)
por: Sun, Hao, et al.
Publicado: (2023)
Temporal Difference Calibration in Sequential Tasks: Application to Vision-Language-Action Models
por: Francis-Meretzki, Shelly, et al.
Publicado: (2026)
por: Francis-Meretzki, Shelly, et al.
Publicado: (2026)
Provably Efficient Exploration in Inverse Constrained Reinforcement Learning
por: Yue, Bo, et al.
Publicado: (2024)
por: Yue, Bo, et al.
Publicado: (2024)
Eidetic Learning: an Efficient and Provable Solution to Catastrophic Forgetting
por: Dronen, Nicholas, et al.
Publicado: (2025)
por: Dronen, Nicholas, et al.
Publicado: (2025)
One-Step Bellman Alignment Enables Provably Efficient Transfer in Online RL
por: Chen, Elynn, et al.
Publicado: (2026)
por: Chen, Elynn, et al.
Publicado: (2026)
Interpetable Target-Feature Aggregation for Multi-Task Learning based on Bias-Variance Analysis
por: Bonetti, Paolo, et al.
Publicado: (2024)
por: Bonetti, Paolo, et al.
Publicado: (2024)
Open Problem: Tight Bounds for Kernelized Multi-Armed Bandits with Bernoulli Rewards
por: Mussi, Marco, et al.
Publicado: (2024)
por: Mussi, Marco, et al.
Publicado: (2024)
Sliding-Window Thompson Sampling for Non-Stationary Settings
por: Fiandri, Marco, et al.
Publicado: (2024)
por: Fiandri, Marco, et al.
Publicado: (2024)
Achieving $\widetilde{\mathcal{O}}(\sqrt{T})$ Regret in Average-Reward POMDPs with Known Observation Models
por: Russo, Alessio, et al.
Publicado: (2025)
por: Russo, Alessio, et al.
Publicado: (2025)
Generalized Kernelized Bandits: A Novel Self-Normalized Bernstein-Like Dimension-Free Inequality and Regret Bounds
por: Metelli, Alberto Maria, et al.
Publicado: (2025)
por: Metelli, Alberto Maria, et al.
Publicado: (2025)
Pure Exploration under Mediators' Feedback
por: Poiani, Riccardo, et al.
Publicado: (2023)
por: Poiani, Riccardo, et al.
Publicado: (2023)
A Refined Analysis of UCBVI
por: Drago, Simone, et al.
Publicado: (2025)
por: Drago, Simone, et al.
Publicado: (2025)
Exploiting Causal Graph Priors with Posterior Sampling for Reinforcement Learning
por: Mutti, Mirco, et al.
Publicado: (2023)
por: Mutti, Mirco, et al.
Publicado: (2023)
Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis
por: Huang, Ruiquan, et al.
Publicado: (2025)
por: Huang, Ruiquan, et al.
Publicado: (2025)
MatRL: Provably Generalizable Iterative Algorithm Discovery via Monte-Carlo Tree Search
por: Kim, Sungyoon, et al.
Publicado: (2025)
por: Kim, Sungyoon, et al.
Publicado: (2025)
An Empirical Risk Minimization Approach for Offline Inverse RL and Dynamic Discrete Choice Model
por: Kang, Enoch H., et al.
Publicado: (2025)
por: Kang, Enoch H., et al.
Publicado: (2025)
Ejemplares similares
-
How does Inverse RL Scale to Large State Spaces? A Provably Efficient Approach
por: Lazzati, Filippo, et al.
Publicado: (2024) -
Reward Compatibility: A Framework for Inverse RL
por: Lazzati, Filippo, et al.
Publicado: (2025) -
Generalizing Behavior via Inverse Reinforcement Learning with Closed-Form Reward Centroids
por: Lazzati, Filippo, et al.
Publicado: (2025) -
Learning Utilities from Demonstrations in Markov Decision Processes
por: Lazzati, Filippo, et al.
Publicado: (2024) -
Robustness in the Face of Partial Identifiability in Reward Learning
por: Lazzati, Filippo, et al.
Publicado: (2025)