Reparameterization Proximal Policy Optimization
Fuente:
arXiv
Saved in:
| Main Authors: | Zhong, Hai, Wang, Xun, Li, Zhuoran, Huang, Longbo |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Reparameterization Flow Policy Optimization
by: Zhong, Hai, et al.
Published: (2026)
by: Zhong, Hai, et al.
Published: (2026)
Beyond Shallow Behavior: Task-Efficient Value-Based Multi-Task Offline MARL via Skill Discovery
by: Wang, Xun, et al.
Published: (2025)
by: Wang, Xun, et al.
Published: (2025)
OM2P: Offline Multi-Agent Mean-Flow Policy
by: Li, Zhuoran, et al.
Published: (2025)
by: Li, Zhuoran, et al.
Published: (2025)
Beyond the Proxy: Trajectory-Distilled Guidance for Offline GFlowNet Training
by: Chen, Ruishuo, et al.
Published: (2025)
by: Chen, Ruishuo, et al.
Published: (2025)
Offline-to-Online Multi-Agent Reinforcement Learning with Offline Value Function Memory and Sequential Exploration
by: Zhong, Hai, et al.
Published: (2024)
by: Zhong, Hai, et al.
Published: (2024)
Beyond Squared Error: Exploring Loss Design for Enhanced Training of Generative Flow Networks
by: Hu, Rui, et al.
Published: (2024)
by: Hu, Rui, et al.
Published: (2024)
Diffusing to Coordinate: Efficient Online Multi-Agent Diffusion Policies
by: Li, Zhuoran, et al.
Published: (2026)
by: Li, Zhuoran, et al.
Published: (2026)
Offline Critic-Guided Diffusion Policy for Multi-User Delay-Constrained Scheduling
by: Li, Zhuoran, et al.
Published: (2025)
by: Li, Zhuoran, et al.
Published: (2025)
PowerFlow: Unlocking the Dual Nature of LLMs via Principled Distribution Matching
by: Chen, Ruishuo, et al.
Published: (2026)
by: Chen, Ruishuo, et al.
Published: (2026)
On-Policy Optimization of ANFIS Policies Using Proximal Policy Optimization
by: Shankar, Kaaustaaub, et al.
Published: (2025)
by: Shankar, Kaaustaaub, et al.
Published: (2025)
ESPO: Early-Stopping Proximal Policy Optimization
by: Li, Zihang, et al.
Published: (2026)
by: Li, Zihang, et al.
Published: (2026)
From Solo to Symphony: Orchestrating Multi-Agent Collaboration with Single-Agent Demos
by: Wang, Xun, et al.
Published: (2025)
by: Wang, Xun, et al.
Published: (2025)
Network Topology Optimization via Deep Reinforcement Learning
by: Li, Zhuoran, et al.
Published: (2022)
by: Li, Zhuoran, et al.
Published: (2022)
Complexity-Regularized Proximal Policy Optimization
by: Serfilippi, Luca, et al.
Published: (2025)
by: Serfilippi, Luca, et al.
Published: (2025)
Beyond the Boundaries of Proximal Policy Optimization
by: Tan, Charlie B., et al.
Published: (2024)
by: Tan, Charlie B., et al.
Published: (2024)
Proximal Policy Optimization with Adaptive Exploration
by: Lixandru, Andrei
Published: (2024)
by: Lixandru, Andrei
Published: (2024)
KIPPO: Koopman-Inspired Proximal Policy Optimization
by: Cozma, Andrei, et al.
Published: (2025)
by: Cozma, Andrei, et al.
Published: (2025)
Diffusion Fine-Tuning via Reparameterized Policy Gradient of the Soft Q-Function
by: Kang, Hyeongyu, et al.
Published: (2025)
by: Kang, Hyeongyu, et al.
Published: (2025)
SofT-GRPO: Surpassing Discrete-Token LLM Reinforcement Learning via Gumbel-Reparameterized Soft-Thinking Policy Optimization
by: Zheng, Zhi, et al.
Published: (2025)
by: Zheng, Zhi, et al.
Published: (2025)
Finite-time Convergence Analysis of Actor-Critic with Evolving Reward
by: Hu, Rui, et al.
Published: (2025)
by: Hu, Rui, et al.
Published: (2025)
Layer-Aware Influence for Online Data Valuation Estimation
by: Yang, Ziao, et al.
Published: (2025)
by: Yang, Ziao, et al.
Published: (2025)
Real-Time Parallel Counterfactual Regret Minimization
by: Li, Boning, et al.
Published: (2026)
by: Li, Boning, et al.
Published: (2026)
CIM-PPO:Proximal Policy Optimization with Liu-Correntropy Induced Metric
by: Guo, Yunxiao, et al.
Published: (2021)
by: Guo, Yunxiao, et al.
Published: (2021)
ExO-PPO: an Extended Off-policy Proximal Policy Optimization Algorithm
by: Wang, Hanyong, et al.
Published: (2026)
by: Wang, Hanyong, et al.
Published: (2026)
Asymmetric Proximal Policy Optimization: mini-critics boost LLM reasoning
by: Liu, Jiashun, et al.
Published: (2025)
by: Liu, Jiashun, et al.
Published: (2025)
Learning Branching Policies for MILPs with Proximal Policy Optimization
by: Mhamed, Abdelouahed Ben, et al.
Published: (2025)
by: Mhamed, Abdelouahed Ben, et al.
Published: (2025)
Proximal Policy Distillation
by: Spigler, Giacomo
Published: (2024)
by: Spigler, Giacomo
Published: (2024)
Theoretical Analysis of Sparse Optimization with Reparameterization, Weight Decay, and Adaptive Learning Rate
by: Xu, Huangyu, et al.
Published: (2026)
by: Xu, Huangyu, et al.
Published: (2026)
A dynamical clipping approach with task feedback for Proximal Policy Optimization
by: Zhang, Ziqi, et al.
Published: (2023)
by: Zhang, Ziqi, et al.
Published: (2023)
Efficient Deep Reinforcement Learning with Predictive Processing Proximal Policy Optimization
by: Küçükoğlu, Burcu, et al.
Published: (2022)
by: Küçükoğlu, Burcu, et al.
Published: (2022)
Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards
by: Ahmad, Ahmad, et al.
Published: (2024)
by: Ahmad, Ahmad, et al.
Published: (2024)
A Deep Reinforcement Learning Approach to Battery Management in Dairy Farming via Proximal Policy Optimization
by: Ali, Nawazish, et al.
Published: (2024)
by: Ali, Nawazish, et al.
Published: (2024)
$\clubsuit$ CLOVER $\clubsuit$: Probabilistic Forecasting with Coherent Learning Objective Reparameterization
by: Olivares, Kin G., et al.
Published: (2023)
by: Olivares, Kin G., et al.
Published: (2023)
ReMAP: Neural Reparameterization for Scalable MAP Inference in Arbitrary-Order Markov Random Fields
by: Wang, Yaomin, et al.
Published: (2024)
by: Wang, Yaomin, et al.
Published: (2024)
AM-PPO: (Advantage) Alpha-Modulation with Proximal Policy Optimization
by: Sane, Soham
Published: (2025)
by: Sane, Soham
Published: (2025)
Proximal Policy Gradient Arborescence for Quality Diversity Reinforcement Learning
by: Batra, Sumeet, et al.
Published: (2023)
by: Batra, Sumeet, et al.
Published: (2023)
Proximity Matters: Local Proximity Enhanced Balancing for Treatment Effect Estimation
by: Wang, Hao, et al.
Published: (2024)
by: Wang, Hao, et al.
Published: (2024)
Variational Delayed Policy Optimization
by: Wu, Qingyuan, et al.
Published: (2024)
by: Wu, Qingyuan, et al.
Published: (2024)
Trust-Region Adaptive Policy Optimization
by: Su, Mingyu, et al.
Published: (2025)
by: Su, Mingyu, et al.
Published: (2025)
Orthogonalized Policy Optimization:Policy Optimization as Orthogonal Projection in Hilbert Space
by: Zixian, Wang
Published: (2026)
by: Zixian, Wang
Published: (2026)
Similar Items
-
Reparameterization Flow Policy Optimization
by: Zhong, Hai, et al.
Published: (2026) -
Beyond Shallow Behavior: Task-Efficient Value-Based Multi-Task Offline MARL via Skill Discovery
by: Wang, Xun, et al.
Published: (2025) -
OM2P: Offline Multi-Agent Mean-Flow Policy
by: Li, Zhuoran, et al.
Published: (2025) -
Beyond the Proxy: Trajectory-Distilled Guidance for Offline GFlowNet Training
by: Chen, Ruishuo, et al.
Published: (2025) -
Offline-to-Online Multi-Agent Reinforcement Learning with Offline Value Function Memory and Sequential Exploration
by: Zhong, Hai, et al.
Published: (2024)