Test-Time Regret Minimization in Meta Reinforcement Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Mutti, Mirco, Tamar, Aviv |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Classification View on Meta Learning Bandits
by: Mutti, Mirco, et al.
Published: (2025)
by: Mutti, Mirco, et al.
Published: (2025)
Temporal Difference Calibration in Sequential Tasks: Application to Vision-Language-Action Models
by: Francis-Meretzki, Shelly, et al.
Published: (2026)
by: Francis-Meretzki, Shelly, et al.
Published: (2026)
Blindfolded Experts Generalize Better: Insights from Robotic Manipulation and Videogames
by: Zisselman, Ev, et al.
Published: (2025)
by: Zisselman, Ev, et al.
Published: (2025)
Meta Reinforcement Learning with Finite Training Tasks -- a Density Estimation Approach
by: Rimon, Zohar, et al.
Published: (2022)
by: Rimon, Zohar, et al.
Published: (2022)
Towards Principled Unsupervised Multi-Agent Reinforcement Learning
by: Zamboni, Riccardo, et al.
Published: (2025)
by: Zamboni, Riccardo, et al.
Published: (2025)
MAMBA: an Effective World Model Approach for Meta-Reinforcement Learning
by: Rimon, Zohar, et al.
Published: (2024)
by: Rimon, Zohar, et al.
Published: (2024)
Exploiting Causal Graph Priors with Posterior Sampling for Reinforcement Learning
by: Mutti, Mirco, et al.
Published: (2023)
by: Mutti, Mirco, et al.
Published: (2023)
TGRL: An Algorithm for Teacher Guided Reinforcement Learning
by: Shenfeld, Idan, et al.
Published: (2023)
by: Shenfeld, Idan, et al.
Published: (2023)
Entity-Centric Reinforcement Learning for Object Manipulation from Pixels
by: Haramati, Dan, et al.
Published: (2024)
by: Haramati, Dan, et al.
Published: (2024)
State Entropy Regularization for Robust Reinforcement Learning
by: Ashlag, Yonatan, et al.
Published: (2025)
by: Ashlag, Yonatan, et al.
Published: (2025)
Reward Compatibility: A Framework for Inverse RL
by: Lazzati, Filippo, et al.
Published: (2025)
by: Lazzati, Filippo, et al.
Published: (2025)
Meta-Learning in Self-Play Regret Minimization
by: Sychrovský, David, et al.
Published: (2025)
by: Sychrovský, David, et al.
Published: (2025)
Offline Inverse RL: New Solution Concepts and Provably Efficient Algorithms
by: Lazzati, Filippo, et al.
Published: (2024)
by: Lazzati, Filippo, et al.
Published: (2024)
How does Inverse RL Scale to Large State Spaces? A Provably Efficient Approach
by: Lazzati, Filippo, et al.
Published: (2024)
by: Lazzati, Filippo, et al.
Published: (2024)
DDLP: Unsupervised Object-Centric Video Prediction with Deep Dynamic Latent Particles
by: Daniel, Tal, et al.
Published: (2023)
by: Daniel, Tal, et al.
Published: (2023)
The Limits of Pure Exploration in POMDPs: When the Observation Entropy is Enough
by: Zamboni, Riccardo, et al.
Published: (2024)
by: Zamboni, Riccardo, et al.
Published: (2024)
Warm-up Free Policy Optimization: Improved Regret in Linear Markov Decision Processes
by: Cassel, Asaf, et al.
Published: (2024)
by: Cassel, Asaf, et al.
Published: (2024)
Multi Task Inverse Reinforcement Learning for Common Sense Reward
by: Glazer, Neta, et al.
Published: (2024)
by: Glazer, Neta, et al.
Published: (2024)
Enhancing Diversity in Parallel Agents: A Maximum State Entropy Exploration Story
by: De Paola, Vincenzo, et al.
Published: (2025)
by: De Paola, Vincenzo, et al.
Published: (2025)
K-Myriad: Jump-starting reinforcement learning with unsupervised parallel agents
by: De Paola, Vincenzo, et al.
Published: (2026)
by: De Paola, Vincenzo, et al.
Published: (2026)
RoboArm-NMP: a Learning Environment for Neural Motion Planning
by: Jurgenson, Tom, et al.
Published: (2024)
by: Jurgenson, Tom, et al.
Published: (2024)
A Theoretical Framework for Partially Observed Reward-States in RLHF
by: Kausik, Chinmaya, et al.
Published: (2024)
by: Kausik, Chinmaya, et al.
Published: (2024)
How to Explore with Belief: State Entropy Maximization in POMDPs
by: Zamboni, Riccardo, et al.
Published: (2024)
by: Zamboni, Riccardo, et al.
Published: (2024)
From Parameters to Behaviors: Unsupervised Compression of the Policy Space
by: Tenedini, Davide, et al.
Published: (2025)
by: Tenedini, Davide, et al.
Published: (2025)
Constrained Meta Reinforcement Learning with Provable Test-Time Safety
by: Ni, Tingting, et al.
Published: (2026)
by: Ni, Tingting, et al.
Published: (2026)
Unsupervised Behavioral Compression: Learning Low-Dimensional Policy Manifolds through State-Occupancy Matching
by: Fraschini, Andrea, et al.
Published: (2026)
by: Fraschini, Andrea, et al.
Published: (2026)
Near-Optimal Regret in Linear MDPs with Aggregate Bandit Feedback
by: Cassel, Asaf, et al.
Published: (2024)
by: Cassel, Asaf, et al.
Published: (2024)
Hierarchical Entity-centric Reinforcement Learning with Factored Subgoal Diffusion
by: Haramati, Dan, et al.
Published: (2026)
by: Haramati, Dan, et al.
Published: (2026)
Toward Artificial Palpation: Representation Learning of Touch on Soft Bodies
by: Rimon, Zohar, et al.
Published: (2025)
by: Rimon, Zohar, et al.
Published: (2025)
Explore to Generalize in Zero-Shot RL
by: Zisselman, Ev, et al.
Published: (2023)
by: Zisselman, Ev, et al.
Published: (2023)
Near-Optimal Regret for Policy Optimization in Contextual MDPs with General Offline Function Approximation
by: Levy, Orin, et al.
Published: (2026)
by: Levy, Orin, et al.
Published: (2026)
No-Regret is not enough! Bandits with General Constraints through Adaptive Regret Minimization
by: Bernasconi, Martino, et al.
Published: (2024)
by: Bernasconi, Martino, et al.
Published: (2024)
Bayesian Regret Minimization in Offline Bandits
by: Petrik, Marek, et al.
Published: (2023)
by: Petrik, Marek, et al.
Published: (2023)
Hierarchical Deep Counterfactual Regret Minimization
by: Chen, Jiayu, et al.
Published: (2023)
by: Chen, Jiayu, et al.
Published: (2023)
Invariance-Based Dynamic Regret Minimization
by: Lazzaretto, Margherita, et al.
Published: (2026)
by: Lazzaretto, Margherita, et al.
Published: (2026)
GradMetaNet: An Equivariant Architecture for Learning on Gradients
by: Gelberg, Yoav, et al.
Published: (2025)
by: Gelberg, Yoav, et al.
Published: (2025)
Stagewise Reinforcement Learning and the Geometry of the Regret Landscape
by: Elliott, Chris, et al.
Published: (2026)
by: Elliott, Chris, et al.
Published: (2026)
Geometric Active Exploration in Markov Decision Processes: the Benefit of Abstraction
by: De Santi, Riccardo, et al.
Published: (2024)
by: De Santi, Riccardo, et al.
Published: (2024)
Heterogeneous Knowledge for Augmented Modular Reinforcement Learning
by: Wolf, Lorenz, et al.
Published: (2023)
by: Wolf, Lorenz, et al.
Published: (2023)
Real-Time Parallel Counterfactual Regret Minimization
by: Li, Boning, et al.
Published: (2026)
by: Li, Boning, et al.
Published: (2026)
Similar Items
-
A Classification View on Meta Learning Bandits
by: Mutti, Mirco, et al.
Published: (2025) -
Temporal Difference Calibration in Sequential Tasks: Application to Vision-Language-Action Models
by: Francis-Meretzki, Shelly, et al.
Published: (2026) -
Blindfolded Experts Generalize Better: Insights from Robotic Manipulation and Videogames
by: Zisselman, Ev, et al.
Published: (2025) -
Meta Reinforcement Learning with Finite Training Tasks -- a Density Estimation Approach
by: Rimon, Zohar, et al.
Published: (2022) -
Towards Principled Unsupervised Multi-Agent Reinforcement Learning
by: Zamboni, Riccardo, et al.
Published: (2025)