Policy Optimization in a Noisy Neighborhood: On Return Landscapes in Continuous Control
Fuente:
arXiv
Saved in:
| Main Authors: | Rahn, Nate, D'Oro, Pierluca, Wiltzer, Harley, Bacon, Pierre-Luc, Bellemare, Marc G. |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Controlling Large Language Model Agents with Entropic Activation Steering
by: Rahn, Nate, et al.
Published: (2024)
by: Rahn, Nate, et al.
Published: (2024)
Mol-MoE: Training Preference-Guided Routers for Molecule Generation
by: Calanzone, Diego, et al.
Published: (2025)
by: Calanzone, Diego, et al.
Published: (2025)
Do Transformer World Models Give Better Policy Gradients?
by: Ma, Michel, et al.
Published: (2024)
by: Ma, Michel, et al.
Published: (2024)
Action Gaps and Advantages in Continuous-Time Distributional Reinforcement Learning
by: Wiltzer, Harley, et al.
Published: (2024)
by: Wiltzer, Harley, et al.
Published: (2024)
Convergence Theorems for Entropy-Regularized and Distributional Reinforcement Learning
by: Jhaveri, Yash, et al.
Published: (2025)
by: Jhaveri, Yash, et al.
Published: (2025)
Maxwell's Demon at Work: Efficient Pruning by Leveraging Saturation of Neurons
by: Dufort-Labbé, Simon, et al.
Published: (2024)
by: Dufort-Labbé, Simon, et al.
Published: (2024)
The Curse of Diversity in Ensemble-Based Exploration
by: Lin, Zhixuan, et al.
Published: (2024)
by: Lin, Zhixuan, et al.
Published: (2024)
ADEPTS: A Capability Framework for Human-Centered Agent Design
by: D'Oro, Pierluca, et al.
Published: (2025)
by: D'Oro, Pierluca, et al.
Published: (2025)
A Distributional Analogue to the Successor Representation
by: Wiltzer, Harley, et al.
Published: (2024)
by: Wiltzer, Harley, et al.
Published: (2024)
Towards General-Purpose Model-Free Reinforcement Learning
by: Fujimoto, Scott, et al.
Published: (2025)
by: Fujimoto, Scott, et al.
Published: (2025)
MaestroMotif: Skill Design from Artificial Intelligence Feedback
by: Klissarov, Martin, et al.
Published: (2024)
by: Klissarov, Martin, et al.
Published: (2024)
Tractable Representations for Convergent Approximation of Distributional HJB Equations
by: Alhosh, Julie, et al.
Published: (2025)
by: Alhosh, Julie, et al.
Published: (2025)
Controlling Multimodal LLMs via Reward-guided Decoding
by: Mañas, Oscar, et al.
Published: (2025)
by: Mañas, Oscar, et al.
Published: (2025)
Hierarchical Behaviour Spaces
by: Matthews, Michael Tryfan, et al.
Published: (2026)
by: Matthews, Michael Tryfan, et al.
Published: (2026)
Foundations of Multivariate Distributional Reinforcement Learning
by: Wiltzer, Harley, et al.
Published: (2024)
by: Wiltzer, Harley, et al.
Published: (2024)
Decoupling regularization from the action space
by: Mohammadpour, Sobhan, et al.
Published: (2024)
by: Mohammadpour, Sobhan, et al.
Published: (2024)
KerJEPA: Kernel Discrepancies for Euclidean Self-Supervised Learning
by: Zimmermann, Eric, et al.
Published: (2025)
by: Zimmermann, Eric, et al.
Published: (2025)
Abstractive Red-Teaming of Language Model Character
by: Rahn, Nate, et al.
Published: (2026)
by: Rahn, Nate, et al.
Published: (2026)
Non-Adversarial Inverse Reinforcement Learning via Successor Feature Matching
by: Jain, Arnav Kumar, et al.
Published: (2024)
by: Jain, Arnav Kumar, et al.
Published: (2024)
Maximum entropy GFlowNets with soft Q-learning
by: Mohammadpour, Sobhan, et al.
Published: (2023)
by: Mohammadpour, Sobhan, et al.
Published: (2023)
On the geometry and topology of representations: the manifolds of modular addition
by: Moisescu-Pareja, Gabriela, et al.
Published: (2025)
by: Moisescu-Pareja, Gabriela, et al.
Published: (2025)
Layerwise LQR for Geometry-Aware Optimization of Deep Networks
by: Dufort-Labbé, Simon, et al.
Published: (2026)
by: Dufort-Labbé, Simon, et al.
Published: (2026)
Fractal Landscapes in Policy Optimization
by: Wang, Tao, et al.
Published: (2023)
by: Wang, Tao, et al.
Published: (2023)
Discovery of Sustainable Refrigerants through Physics-Informed RL Fine-Tuning of Sequence Models
by: Goldszal, Adrien, et al.
Published: (2025)
by: Goldszal, Adrien, et al.
Published: (2025)
The Three Regimes of Offline-to-Online Reinforcement Learning
by: Li, Lu, et al.
Published: (2025)
by: Li, Lu, et al.
Published: (2025)
ANO : Faster is Better in Noisy Landscape
by: Kegreisz, Adrien
Published: (2025)
by: Kegreisz, Adrien
Published: (2025)
Tapered Off-Policy REINFORCE: Stable and efficient reinforcement learning for LLMs
by: Roux, Nicolas Le, et al.
Published: (2025)
by: Roux, Nicolas Le, et al.
Published: (2025)
On Symmetric Losses for Robust Policy Optimization with Noisy Preferences
by: Nishimori, Soichiro, et al.
Published: (2025)
by: Nishimori, Soichiro, et al.
Published: (2025)
Moments Matter:Stabilizing Policy Optimization using Return Distributions
by: Jabs, Dennis, et al.
Published: (2026)
by: Jabs, Dennis, et al.
Published: (2026)
Compositional Planning with Jumpy World Models
by: Farebrother, Jesse, et al.
Published: (2026)
by: Farebrother, Jesse, et al.
Published: (2026)
Understanding Behavioral Metric Learning: A Large-Scale Study on Distracting Reinforcement Learning Environments
by: Luo, Ziyan, et al.
Published: (2025)
by: Luo, Ziyan, et al.
Published: (2025)
Reward Redistribution for CVaR MDPs using a Bellman Operator on L-infinity
by: Muni, Aneri, et al.
Published: (2026)
by: Muni, Aneri, et al.
Published: (2026)
Long-Horizon Model-Based Offline Reinforcement Learning Without Explicit Conservatism
by: Ni, Tianwei, et al.
Published: (2025)
by: Ni, Tianwei, et al.
Published: (2025)
State Entropy Regularization for Robust Reinforcement Learning
by: Ashlag, Yonatan, et al.
Published: (2025)
by: Ashlag, Yonatan, et al.
Published: (2025)
What Makes Value Learning Efficient in Residual Reinforcement Learning?
by: Ma, Guozheng, et al.
Published: (2026)
by: Ma, Guozheng, et al.
Published: (2026)
Multitask Online Learning: Listen to the Neighborhood Buzz
by: Achddou, Juliette, et al.
Published: (2023)
by: Achddou, Juliette, et al.
Published: (2023)
Kinematic Tokenization: Optimization-Based Continuous-Time Tokens for Learnable Decision Policies in Noisy Time Series
by: Kearney, Griffin
Published: (2026)
by: Kearney, Griffin
Published: (2026)
Behaviour Policy Optimization: Provably Lower Variance Return Estimates for Off-Policy Reinforcement Learning
by: Goodall, Alexander W., et al.
Published: (2025)
by: Goodall, Alexander W., et al.
Published: (2025)
SFBD Flow: A Continuous-Optimization Framework for Training Diffusion Models with Noisy Samples
by: Lu, Haoye, et al.
Published: (2025)
by: Lu, Haoye, et al.
Published: (2025)
Boosting Continuous Control with Consistency Policy
by: Chen, Yuhui, et al.
Published: (2023)
by: Chen, Yuhui, et al.
Published: (2023)
Similar Items
-
Controlling Large Language Model Agents with Entropic Activation Steering
by: Rahn, Nate, et al.
Published: (2024) -
Mol-MoE: Training Preference-Guided Routers for Molecule Generation
by: Calanzone, Diego, et al.
Published: (2025) -
Do Transformer World Models Give Better Policy Gradients?
by: Ma, Michel, et al.
Published: (2024) -
Action Gaps and Advantages in Continuous-Time Distributional Reinforcement Learning
by: Wiltzer, Harley, et al.
Published: (2024) -
Convergence Theorems for Entropy-Regularized and Distributional Reinforcement Learning
by: Jhaveri, Yash, et al.
Published: (2025)