Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying
Fuente:
arXiv
Saved in:
| Main Authors: | Nishimori, Soichiro, Parmas, Paavo, Koyamada, Sotetsu, Kozuno, Tadashi, Kitamura, Toshinori, Ishii, Shin, Matsuo, Yutaka |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Batch Sequential Halving Algorithm without Performance Degradation
by: Koyamada, Sotetsu, et al.
Published: (2024)
by: Koyamada, Sotetsu, et al.
Published: (2024)
Pgx: Hardware-Accelerated Parallel Game Simulators for Reinforcement Learning
by: Koyamada, Sotetsu, et al.
Published: (2023)
by: Koyamada, Sotetsu, et al.
Published: (2023)
Does "Do Differentiable Simulators Give Better Policy Gradients?'' Give Better Policy Gradients?
by: Onoda, Ku, et al.
Published: (2026)
by: Onoda, Ku, et al.
Published: (2026)
Finite-Time Regret Analysis of Retry-Aware Bandits
by: Tong, Bingkui, et al.
Published: (2026)
by: Tong, Bingkui, et al.
Published: (2026)
Mahjax: A GPU-Accelerated Mahjong Simulator for Reinforcement Learning in JAX
by: Nishimori, Soichiro, et al.
Published: (2026)
by: Nishimori, Soichiro, et al.
Published: (2026)
Double Horizon Model-Based Policy Optimization
by: Kubo, Akihiro, et al.
Published: (2025)
by: Kubo, Akihiro, et al.
Published: (2025)
Near-Optimal Policy Identification in Robust Constrained Markov Decision Processes via Epigraph Form
by: Kitamura, Toshinori, et al.
Published: (2024)
by: Kitamura, Toshinori, et al.
Published: (2024)
A Policy Gradient Primal-Dual Algorithm for Constrained MDPs with Uniform PAC Guarantees
by: Kitamura, Toshinori, et al.
Published: (2024)
by: Kitamura, Toshinori, et al.
Published: (2024)
A Simple, Solid, and Reproducible Baseline for Bridge Bidding AI
by: Kita, Haruka, et al.
Published: (2024)
by: Kita, Haruka, et al.
Published: (2024)
Symmetry-aware Reinforcement Learning for Robotic Assembly under Partial Observability with a Soft Wrist
by: Nguyen, Hai, et al.
Published: (2024)
by: Nguyen, Hai, et al.
Published: (2024)
On Symmetric Losses for Robust Policy Optimization with Noisy Preferences
by: Nishimori, Soichiro, et al.
Published: (2025)
by: Nishimori, Soichiro, et al.
Published: (2025)
R$^3$L: Reflect-then-Retry Reinforcement Learning with Language-Guided Exploration, Pivotal Credit, and Positive Amplification
by: Shi, Weijie, et al.
Published: (2026)
by: Shi, Weijie, et al.
Published: (2026)
Provably Efficient RL under Episode-Wise Safety in Constrained MDPs with Linear Function Approximation
by: Kitamura, Toshinori, et al.
Published: (2025)
by: Kitamura, Toshinori, et al.
Published: (2025)
LPPG-RL: Lexicographically Projected Policy Gradient Reinforcement Learning with Subproblem Exploration
by: Qiu, Ruiyu, et al.
Published: (2025)
by: Qiu, Ruiyu, et al.
Published: (2025)
Robust off-policy Reinforcement Learning via Soft Constrained Adversary
by: Nakanishi, Kosuke, et al.
Published: (2024)
by: Nakanishi, Kosuke, et al.
Published: (2024)
Mitigating Reward Hacking in RLHF via Advantage Sign Robustness
by: Ono, Shinnosuke, et al.
Published: (2026)
by: Ono, Shinnosuke, et al.
Published: (2026)
A Comprehensive Survey on Physical Risk Control in the Era of Foundation Model-enabled Robotics
by: Kojima, Takeshi, et al.
Published: (2025)
by: Kojima, Takeshi, et al.
Published: (2025)
Efficient On-Policy Reinforcement Learning via Exploration of Sparse Parameter Space
by: Zhang, Xinyu, et al.
Published: (2025)
by: Zhang, Xinyu, et al.
Published: (2025)
Off-Policy Actor-Critic for Adversarial Observation Robustness: Virtual Alternative Training via Symmetric Policy Evaluation
by: Nakanishi, Kosuke, et al.
Published: (2025)
by: Nakanishi, Kosuke, et al.
Published: (2025)
Policy Gradient Methods for Non-Markovian Reinforcement Learning
by: Kar, Avik, et al.
Published: (2026)
by: Kar, Avik, et al.
Published: (2026)
Policy Gradients for Cumulative Prospect Theory in Reinforcement Learning
by: Lepel, Olivier, et al.
Published: (2024)
by: Lepel, Olivier, et al.
Published: (2024)
Accelerating Reinforcement Learning with Value-Conditional State Entropy Exploration
by: Kim, Dongyoung, et al.
Published: (2023)
by: Kim, Dongyoung, et al.
Published: (2023)
Self-Harmony: Learning to Harmonize Self-Supervision and Self-Play in Test-Time Reinforcement Learning
by: Wang, Ru, et al.
Published: (2025)
by: Wang, Ru, et al.
Published: (2025)
Enhancing Unimodal Latent Representations in Multimodal VAEs through Iterative Amortized Inference
by: Oshima, Yuta, et al.
Published: (2024)
by: Oshima, Yuta, et al.
Published: (2024)
$K$-Level Policy Gradients for Multi-Agent Reinforcement Learning
by: Reddi, Aryaman, et al.
Published: (2025)
by: Reddi, Aryaman, et al.
Published: (2025)
Proximal Policy Gradient Arborescence for Quality Diversity Reinforcement Learning
by: Batra, Sumeet, et al.
Published: (2023)
by: Batra, Sumeet, et al.
Published: (2023)
ProRL: Effective Reinforcement Learning for Proactive Recommendation via Rectified Policy Gradient Estimation
by: Hou, Hongru, et al.
Published: (2026)
by: Hou, Hongru, et al.
Published: (2026)
Can LLMs Guide Their Own Exploration? Gradient-Guided Reinforcement Learning for LLM Reasoning
by: Liang, Zhenwen, et al.
Published: (2025)
by: Liang, Zhenwen, et al.
Published: (2025)
Understanding Emergent Misalignment via Feature Superposition Geometry
by: Minegishi, Gouki, et al.
Published: (2026)
by: Minegishi, Gouki, et al.
Published: (2026)
Rethinking Policy Diversity in Ensemble Policy Gradient in Large-Scale Reinforcement Learning
by: Shitanda, Naoki, et al.
Published: (2026)
by: Shitanda, Naoki, et al.
Published: (2026)
Stabilizing Policy Gradients for Sample-Efficient Reinforcement Learning in LLM Reasoning
by: Melo, Luckeciano C., et al.
Published: (2025)
by: Melo, Luckeciano C., et al.
Published: (2025)
On the Global Optimality of Policy Gradient Methods in General Utility Reinforcement Learning
by: Barakat, Anas, et al.
Published: (2024)
by: Barakat, Anas, et al.
Published: (2024)
PG-Rainbow: Using Distributional Reinforcement Learning in Policy Gradient Methods
by: Jeon, WooJae, et al.
Published: (2024)
by: Jeon, WooJae, et al.
Published: (2024)
Diffusion Fine-Tuning via Reparameterized Policy Gradient of the Soft Q-Function
by: Kang, Hyeongyu, et al.
Published: (2025)
by: Kang, Hyeongyu, et al.
Published: (2025)
Reinforcement Learning via Value Gradient Flow
by: Xu, Haoran, et al.
Published: (2026)
by: Xu, Haoran, et al.
Published: (2026)
Performative Policy Gradient: Optimality in Performative Reinforcement Learning
by: Basu, Debabrota, et al.
Published: (2025)
by: Basu, Debabrota, et al.
Published: (2025)
When Object-Centric World Models Meet Policy Learning: From Pixels to Policies, and Where It Breaks
by: Ferraro, Stefano, et al.
Published: (2025)
by: Ferraro, Stefano, et al.
Published: (2025)
The Definitive Guide to Policy Gradients in Deep Reinforcement Learning: Theory, Algorithms and Implementations
by: Lehmann, Matthias
Published: (2024)
by: Lehmann, Matthias
Published: (2024)
Language Models Do Hard Arithmetic Tasks Easily and Hardly Do Easy Arithmetic Tasks
by: Gambardella, Andrew, et al.
Published: (2024)
by: Gambardella, Andrew, et al.
Published: (2024)
Towards Empirical Interpretation of Internal Circuits and Properties in Grokked Transformers on Modular Polynomials
by: Furuta, Hiroki, et al.
Published: (2024)
by: Furuta, Hiroki, et al.
Published: (2024)
Similar Items
-
A Batch Sequential Halving Algorithm without Performance Degradation
by: Koyamada, Sotetsu, et al.
Published: (2024) -
Pgx: Hardware-Accelerated Parallel Game Simulators for Reinforcement Learning
by: Koyamada, Sotetsu, et al.
Published: (2023) -
Does "Do Differentiable Simulators Give Better Policy Gradients?'' Give Better Policy Gradients?
by: Onoda, Ku, et al.
Published: (2026) -
Finite-Time Regret Analysis of Retry-Aware Bandits
by: Tong, Bingkui, et al.
Published: (2026) -
Mahjax: A GPU-Accelerated Mahjong Simulator for Reinforcement Learning in JAX
by: Nishimori, Soichiro, et al.
Published: (2026)