Bad Habits: Policy Confounding and Out-of-Trajectory Generalization in RL
Fuente:
arXiv
Guardado en:
| Autores principales: | Suau, Miguel, Spaan, Matthijs T. J., Oliehoek, Frans A. |
|---|---|
| Formato: | Preprint |
| Publicado: |
2023
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Distributed Influence-Augmented Local Simulators for Parallel MARL in Large Networked Systems
por: Suau, Miguel, et al.
Publicado: (2022)
por: Suau, Miguel, et al.
Publicado: (2022)
When Do Off-Policy and On-Policy Policy Gradient Methods Align?
por: Mambelli, Davide, et al.
Publicado: (2024)
por: Mambelli, Davide, et al.
Publicado: (2024)
Conditional Policy Generator for Dynamic Constraint Satisfaction and Optimization
por: Lee, Wook, et al.
Publicado: (2025)
por: Lee, Wook, et al.
Publicado: (2025)
Breaking Habits: On the Role of the Advantage Function in Learning Causal State Representations
por: Suau, Miguel
Publicado: (2025)
por: Suau, Miguel
Publicado: (2025)
Explaining Learned Reward Functions with Counterfactual Trajectories
por: Wehner, Jan, et al.
Publicado: (2024)
por: Wehner, Jan, et al.
Publicado: (2024)
Sparse Masked Attention Policies for Reliable Generalization
por: Horsch, Caroline, et al.
Publicado: (2026)
por: Horsch, Caroline, et al.
Publicado: (2026)
SimuDICE: Offline Policy Optimization Through World Model Updates and DICE Estimation
por: Brita, Catalin E., et al.
Publicado: (2024)
por: Brita, Catalin E., et al.
Publicado: (2024)
Sample-Efficient Policy Space Response Oracles with Joint Experience Best Response
por: Bighashdel, Ariyan, et al.
Publicado: (2026)
por: Bighashdel, Ariyan, et al.
Publicado: (2026)
Trust-Region Twisted Policy Improvement
por: de Vries, Joery A., et al.
Publicado: (2025)
por: de Vries, Joery A., et al.
Publicado: (2025)
Off-Policy Safe Reinforcement Learning with Constrained Optimistic Exploration
por: Li, Guopeng, et al.
Publicado: (2026)
por: Li, Guopeng, et al.
Publicado: (2026)
How Ensembles of Distilled Policies Improve Generalisation in Reinforcement Learning
por: Weltevrede, Max, et al.
Publicado: (2025)
por: Weltevrede, Max, et al.
Publicado: (2025)
AlphaExploitem: Going Beyond the Nash Equilibrium in Poker by Learning to Exploit Suboptimal Play
por: Murgoci, Vlad, et al.
Publicado: (2026)
por: Murgoci, Vlad, et al.
Publicado: (2026)
Difference Rewards Policy Gradients
por: Castellini, Jacopo, et al.
Publicado: (2020)
por: Castellini, Jacopo, et al.
Publicado: (2020)
CoMI-IRL: Contrastive Multi-Intention Inverse Reinforcement Learning
por: Mone, Antonio, et al.
Publicado: (2026)
por: Mone, Antonio, et al.
Publicado: (2026)
Navigating Trade-offs: Policy Summarization for Multi-Objective Reinforcement Learning
por: Osika, Zuzanna, et al.
Publicado: (2024)
por: Osika, Zuzanna, et al.
Publicado: (2024)
Diverse Projection Ensembles for Distributional Reinforcement Learning
por: Zanger, Moritz A., et al.
Publicado: (2023)
por: Zanger, Moritz A., et al.
Publicado: (2023)
Positive Experience Reflection for Agents in Interactive Text Environments
por: Lippmann, Philip, et al.
Publicado: (2024)
por: Lippmann, Philip, et al.
Publicado: (2024)
Learning to Focus: Prioritizing Informative Histories with Structured Attention Mechanisms in Partially Observable Reinforcement Learning
por: Allegue, Daniel De Dios, et al.
Publicado: (2025)
por: Allegue, Daniel De Dios, et al.
Publicado: (2025)
Epistemic Monte Carlo Tree Search
por: Oren, Yaniv, et al.
Publicado: (2022)
por: Oren, Yaniv, et al.
Publicado: (2022)
Exploration Implies Data Augmentation: Reachability and Generalisation in Contextual MDPs
por: Weltevrede, Max, et al.
Publicado: (2024)
por: Weltevrede, Max, et al.
Publicado: (2024)
Explore-Go: Leveraging Exploration for Generalisation in Deep Reinforcement Learning
por: Weltevrede, Max, et al.
Publicado: (2024)
por: Weltevrede, Max, et al.
Publicado: (2024)
Timing the Match: A Deep Reinforcement Learning Approach for Ride-Hailing and Ride-Pooling Services
por: Bao, Yiman, et al.
Publicado: (2025)
por: Bao, Yiman, et al.
Publicado: (2025)
Exploring Equity of Climate Policies using Multi-Agent Multi-Objective Reinforcement Learning
por: Biswas, Palok, et al.
Publicado: (2025)
por: Biswas, Palok, et al.
Publicado: (2025)
What model does MuZero learn?
por: He, Jinke, et al.
Publicado: (2023)
por: He, Jinke, et al.
Publicado: (2023)
Bayesian Meta-Reinforcement Learning with Laplace Variational Recurrent Networks
por: de Vries, Joery A., et al.
Publicado: (2025)
por: de Vries, Joery A., et al.
Publicado: (2025)
Priors Matter: Addressing Misspecification in Bayesian Deep Q-Learning
por: van der Vaart, Pascal R., et al.
Publicado: (2025)
por: van der Vaart, Pascal R., et al.
Publicado: (2025)
Online Planning in POMDPs with State-Requests
por: Avalos, Raphael, et al.
Publicado: (2024)
por: Avalos, Raphael, et al.
Publicado: (2024)
Scalable Out-of-distribution Robustness in the Presence of Unobserved Confounders
por: Prashant, Parjanya, et al.
Publicado: (2024)
por: Prashant, Parjanya, et al.
Publicado: (2024)
Contextual Similarity Distillation: Ensemble Uncertainties with a Single Model
por: Zanger, Moritz A., et al.
Publicado: (2025)
por: Zanger, Moritz A., et al.
Publicado: (2025)
Multi-Objective Reinforcement Learning for Water Management
por: Osika, Zuzanna, et al.
Publicado: (2025)
por: Osika, Zuzanna, et al.
Publicado: (2025)
PMCTS: Particle Monte Carlo Tree Search for Principled Parallelized Inference Time Scaling
por: Oren, Yaniv, et al.
Publicado: (2026)
por: Oren, Yaniv, et al.
Publicado: (2026)
Good Actions Succeed, Bad Actions Generalize: A Case Study on Why RL Generalizes Better
por: Song, Meng
Publicado: (2025)
por: Song, Meng
Publicado: (2025)
Reinforcement Learning by Guided Safe Exploration
por: Yang, Qisong, et al.
Publicado: (2023)
por: Yang, Qisong, et al.
Publicado: (2023)
Uncoupled Learning of Differential Stackelberg Equilibria with Commitments
por: Loftin, Robert, et al.
Publicado: (2023)
por: Loftin, Robert, et al.
Publicado: (2023)
Inverse Concave-Utility Reinforcement Learning is Inverse Game Theory
por: Çelikok, Mustafa Mert, et al.
Publicado: (2024)
por: Çelikok, Mustafa Mert, et al.
Publicado: (2024)
RecBayes: Recurrent Bayesian Ad Hoc Teamwork in Large Partially Observable Domains
por: Ribeiro, João G., et al.
Publicado: (2025)
por: Ribeiro, João G., et al.
Publicado: (2025)
On the Equivalence of Random Network Distillation, Deep Ensembles, and Bayesian Inference
por: Zanger, Moritz A., et al.
Publicado: (2026)
por: Zanger, Moritz A., et al.
Publicado: (2026)
Twice Sequential Monte Carlo for Tree Search
por: Oren, Yaniv, et al.
Publicado: (2025)
por: Oren, Yaniv, et al.
Publicado: (2025)
OGBench: Benchmarking Offline Goal-Conditioned RL
por: Park, Seohong, et al.
Publicado: (2024)
por: Park, Seohong, et al.
Publicado: (2024)
Quantile-Optimal Policy Learning under Unmeasured Confounding
por: Chen, Zhongren, et al.
Publicado: (2025)
por: Chen, Zhongren, et al.
Publicado: (2025)
Ejemplares similares
-
Distributed Influence-Augmented Local Simulators for Parallel MARL in Large Networked Systems
por: Suau, Miguel, et al.
Publicado: (2022) -
When Do Off-Policy and On-Policy Policy Gradient Methods Align?
por: Mambelli, Davide, et al.
Publicado: (2024) -
Conditional Policy Generator for Dynamic Constraint Satisfaction and Optimization
por: Lee, Wook, et al.
Publicado: (2025) -
Breaking Habits: On the Role of the Advantage Function in Learning Causal State Representations
por: Suau, Miguel
Publicado: (2025) -
Explaining Learned Reward Functions with Counterfactual Trajectories
por: Wehner, Jan, et al.
Publicado: (2024)