Learning to Optimize for Reinforcement Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Lan, Qingfeng, Mahmood, A. Rupam, Yan, Shuicheng, Xu, Zhongwen |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Weight Clipping for Deep Continual and Reinforcement Learning
by: Elsayed, Mohamed, et al.
Published: (2024)
by: Elsayed, Mohamed, et al.
Published: (2024)
Mutual Information Regularized Offline Reinforcement Learning
by: Ma, Xiao, et al.
Published: (2022)
by: Ma, Xiao, et al.
Published: (2022)
More Efficient Randomized Exploration for Reinforcement Learning via Approximate Sampling
by: Ishfaq, Haque, et al.
Published: (2024)
by: Ishfaq, Haque, et al.
Published: (2024)
Streaming Deep Reinforcement Learning Finally Works
by: Elsayed, Mohamed, et al.
Published: (2024)
by: Elsayed, Mohamed, et al.
Published: (2024)
Addressing Loss of Plasticity and Catastrophic Forgetting in Continual Learning
by: Elsayed, Mohamed, et al.
Published: (2024)
by: Elsayed, Mohamed, et al.
Published: (2024)
Efficient Reinforcement Learning by Reducing Forgetting with Elephant Activation Functions
by: Lan, Qingfeng, et al.
Published: (2025)
by: Lan, Qingfeng, et al.
Published: (2025)
Revisiting Scalable Hessian Diagonal Approximations for Applications in Reinforcement Learning
by: Elsayed, Mohamed, et al.
Published: (2024)
by: Elsayed, Mohamed, et al.
Published: (2024)
Intentional Updates for Streaming Reinforcement Learning
by: Sharifnassab, Arsalan, et al.
Published: (2026)
by: Sharifnassab, Arsalan, et al.
Published: (2026)
Reinforcement Learning from Diverse Human Preferences
by: Xue, Wanqi, et al.
Published: (2023)
by: Xue, Wanqi, et al.
Published: (2023)
Single-stream Policy Optimization
by: Xu, Zhongwen, et al.
Published: (2025)
by: Xu, Zhongwen, et al.
Published: (2025)
Policy Regularization on Globally Accessible States in Cross-Dynamics Reinforcement Learning
by: Xue, Zhenghai, et al.
Published: (2025)
by: Xue, Zhenghai, et al.
Published: (2025)
Distributions as Actions: A Unified Framework for Diverse Action Spaces
by: He, Jiamin, et al.
Published: (2025)
by: He, Jiamin, et al.
Published: (2025)
Cogito, Ergo Ludo: An Agent that Learns to Play by Reasoning and Planning
by: Wang, Sai, et al.
Published: (2025)
by: Wang, Sai, et al.
Published: (2025)
Operator-Guided Invariance Learning for Continuous Reinforcement Learning
by: Zhang, Zuyuan, et al.
Published: (2026)
by: Zhang, Zuyuan, et al.
Published: (2026)
Provable and Practical: Efficient Exploration in Reinforcement Learning via Langevin Monte Carlo
by: Ishfaq, Haque, et al.
Published: (2023)
by: Ishfaq, Haque, et al.
Published: (2023)
Understanding Tool-Integrated Reasoning
by: Lin, Heng, et al.
Published: (2025)
by: Lin, Heng, et al.
Published: (2025)
Diffusion Policies with Value-Conditional Optimization for Offline Reinforcement Learning
by: Ma, Yunchang, et al.
Published: (2025)
by: Ma, Yunchang, et al.
Published: (2025)
OffSeeker: Online Reinforcement Learning Is Not All You Need for Deep Research Agents
by: Zhou, Yuhang, et al.
Published: (2026)
by: Zhou, Yuhang, et al.
Published: (2026)
Conformal Symplectic Optimization for Stable Reinforcement Learning
by: Lyu, Yao, et al.
Published: (2024)
by: Lyu, Yao, et al.
Published: (2024)
Percentile Criterion Optimization in Offline Reinforcement Learning
by: Lobo, Elita A., et al.
Published: (2024)
by: Lobo, Elita A., et al.
Published: (2024)
QF-tuner: Breaking Tradition in Reinforcement Learning
by: Jumaah, Mahmood A., et al.
Published: (2024)
by: Jumaah, Mahmood A., et al.
Published: (2024)
Learning to Synthesize Compatible Fashion Items Using Semantic Alignment and Collocation Classification: An Outfit Generation Framework
by: Zhou, Dongliang, et al.
Published: (2025)
by: Zhou, Dongliang, et al.
Published: (2025)
Can Learned Optimization Make Reinforcement Learning Less Difficult?
by: Goldie, Alexander David, et al.
Published: (2024)
by: Goldie, Alexander David, et al.
Published: (2024)
DUET: Optimize Token-Budget Allocation for Reinforcement Learning with Verifiable Rewards
by: Hu, Haoyu, et al.
Published: (2026)
by: Hu, Haoyu, et al.
Published: (2026)
Matrix-Space Reinforcement Learning for Reusing Local Transition Geometry
by: Zhang, Zuyuan, et al.
Published: (2026)
by: Zhang, Zuyuan, et al.
Published: (2026)
GEPO: Group Expectation Policy Optimization for Stable Heterogeneous Reinforcement Learning
by: Zhang, Han, et al.
Published: (2025)
by: Zhang, Han, et al.
Published: (2025)
Online Optimization for Offline Safe Reinforcement Learning
by: Chemingui, Yassine, et al.
Published: (2025)
by: Chemingui, Yassine, et al.
Published: (2025)
Optimizing Automatic Differentiation with Deep Reinforcement Learning
by: Lohoff, Jamie, et al.
Published: (2024)
by: Lohoff, Jamie, et al.
Published: (2024)
Performance Optimization of Ratings-Based Reinforcement Learning
by: Rose, Evelyn, et al.
Published: (2025)
by: Rose, Evelyn, et al.
Published: (2025)
Offline Trajectory Optimization for Offline Reinforcement Learning
by: Zhao, Ziqi, et al.
Published: (2024)
by: Zhao, Ziqi, et al.
Published: (2024)
Is Exploration or Optimization the Problem for Deep Reinforcement Learning?
by: Berseth, Glen
Published: (2025)
by: Berseth, Glen
Published: (2025)
MoE++: Accelerating Mixture-of-Experts Methods with Zero-Computation Experts
by: Jin, Peng, et al.
Published: (2024)
by: Jin, Peng, et al.
Published: (2024)
Adaptive Advantage-Guided Policy Regularization for Offline Reinforcement Learning
by: Liu, Tenglong, et al.
Published: (2024)
by: Liu, Tenglong, et al.
Published: (2024)
Manifold-Constrained Energy-Based Transition Models for Offline Reinforcement Learning
by: Fang, Zeyu, et al.
Published: (2026)
by: Fang, Zeyu, et al.
Published: (2026)
ODRL: A Benchmark for Off-Dynamics Reinforcement Learning
by: Lyu, Jiafei, et al.
Published: (2024)
by: Lyu, Jiafei, et al.
Published: (2024)
Quantum Circuit Structure Optimization for Quantum Reinforcement Learning
by: Son, Seok Bin, et al.
Published: (2025)
by: Son, Seok Bin, et al.
Published: (2025)
Applying Reinforcement Learning to Optimize Traffic Light Cycles
by: Son, Seungah, et al.
Published: (2024)
by: Son, Seungah, et al.
Published: (2024)
Robust Federated Finetuning of LLMs via Alternating Optimization of LoRA
by: Chen, Shuangyi, et al.
Published: (2025)
by: Chen, Shuangyi, et al.
Published: (2025)
RetICL: Sequential Retrieval of In-Context Examples with Reinforcement Learning
by: Scarlatos, Alexander, et al.
Published: (2023)
by: Scarlatos, Alexander, et al.
Published: (2023)
Constraint-Conditioned Policy Optimization for Versatile Safe Reinforcement Learning
by: Yao, Yihang, et al.
Published: (2023)
by: Yao, Yihang, et al.
Published: (2023)
Similar Items
-
Weight Clipping for Deep Continual and Reinforcement Learning
by: Elsayed, Mohamed, et al.
Published: (2024) -
Mutual Information Regularized Offline Reinforcement Learning
by: Ma, Xiao, et al.
Published: (2022) -
More Efficient Randomized Exploration for Reinforcement Learning via Approximate Sampling
by: Ishfaq, Haque, et al.
Published: (2024) -
Streaming Deep Reinforcement Learning Finally Works
by: Elsayed, Mohamed, et al.
Published: (2024) -
Addressing Loss of Plasticity and Catastrophic Forgetting in Continual Learning
by: Elsayed, Mohamed, et al.
Published: (2024)