Saved in:
| Main Authors: | Zhang, Ruipeng, Chang, Ya-Chien, Gao, Sicun |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2506.05615 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Extremum-Seeking Action Selection for Accelerating Policy Optimization
by: Chang, Ya-Chien, et al.
Published: (2024)
by: Chang, Ya-Chien, et al.
Published: (2024)
Improving Value Estimation Critically Enhances Vanilla Policy Gradient
by: Wang, Tao, et al.
Published: (2025)
by: Wang, Tao, et al.
Published: (2025)
Learning Quadruped Walking from Seconds of Demonstration
by: Zhang, Ruipeng, et al.
Published: (2026)
by: Zhang, Ruipeng, et al.
Published: (2026)
Fractal Landscapes in Policy Optimization
by: Wang, Tao, et al.
Published: (2023)
by: Wang, Tao, et al.
Published: (2023)
Activation-Descent Regularization for Input Optimization of ReLU Networks
by: Yu, Hongzhan, et al.
Published: (2024)
by: Yu, Hongzhan, et al.
Published: (2024)
Mollification Effects of Policy Gradient Methods
by: Wang, Tao, et al.
Published: (2024)
by: Wang, Tao, et al.
Published: (2024)
Maximum Entropy Reinforcement Learning with Diffusion Policy
by: Dong, Xiaoyi, et al.
Published: (2025)
by: Dong, Xiaoyi, et al.
Published: (2025)
Improving Compositional Generation with Diffusion Models Using Lift Scores
by: Yu, Chenning, et al.
Published: (2025)
by: Yu, Chenning, et al.
Published: (2025)
Maximum Entropy On-Policy Actor-Critic via Entropy Advantage Estimation
by: Choe, Jean Seong Bjorn, et al.
Published: (2024)
by: Choe, Jean Seong Bjorn, et al.
Published: (2024)
Breaking the Barrier: Enhanced Utility and Robustness in Smoothed DRL Agents
by: Sun, Chung-En, et al.
Published: (2024)
by: Sun, Chung-En, et al.
Published: (2024)
ESPO: Entropy Importance Sampling Policy Optimization
by: Sheng, Yuepeng, et al.
Published: (2025)
by: Sheng, Yuepeng, et al.
Published: (2025)
Mind Your Entropy: From Maximum Entropy to Trajectory Entropy-Constrained RL
by: Zhan, Guojian, et al.
Published: (2025)
by: Zhan, Guojian, et al.
Published: (2025)
Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization
by: Xu, Huimin, et al.
Published: (2026)
by: Xu, Huimin, et al.
Published: (2026)
Maximum Entropy Exploration Without the Rollouts
by: Adamczyk, Jacob, et al.
Published: (2026)
by: Adamczyk, Jacob, et al.
Published: (2026)
Entropy-Regularized Token-Level Policy Optimization for Language Agent Reinforcement
by: Wen, Muning, et al.
Published: (2024)
by: Wen, Muning, et al.
Published: (2024)
Deriving the Scaled-Dot-Function via Maximum Likelihood Estimation and Maximum Entropy Approach
by: Ma, Jiyong
Published: (2025)
by: Ma, Jiyong
Published: (2025)
Quagmires in SFT-RL Post-Training: When High SFT Scores Mislead and What to Use Instead
by: Kang, Feiyang, et al.
Published: (2025)
by: Kang, Feiyang, et al.
Published: (2025)
When are LLMs Sufficient Policy Optimizers for Sequential RL Tasks?
by: Hatgis-Kessell, Stephane, et al.
Published: (2026)
by: Hatgis-Kessell, Stephane, et al.
Published: (2026)
ELEMENT: Episodic and Lifelong Exploration via Maximum Entropy
by: Li, Hongming, et al.
Published: (2024)
by: Li, Hongming, et al.
Published: (2024)
Reward-Punishment Reinforcement Learning with Maximum Entropy
by: Wang, Jiexin, et al.
Published: (2024)
by: Wang, Jiexin, et al.
Published: (2024)
Agentic Entropy-Balanced Policy Optimization
by: Dong, Guanting, et al.
Published: (2025)
by: Dong, Guanting, et al.
Published: (2025)
FastDSAC: Unlocking the Potential of Maximum Entropy RL in High-Dimensional Humanoid Control
by: Xue, Jun, et al.
Published: (2026)
by: Xue, Jun, et al.
Published: (2026)
Learning-based Motion Planning in Dynamic Environments Using GNNs and Temporal Encoding
by: Zhang, Ruipeng, et al.
Published: (2022)
by: Zhang, Ruipeng, et al.
Published: (2022)
ERPO: Token-Level Entropy-Regulated Policy Optimization for Large Reasoning Models
by: Yu, Song, et al.
Published: (2026)
by: Yu, Song, et al.
Published: (2026)
Fibration Policy Optimization
by: Li, Chang, et al.
Published: (2026)
by: Li, Chang, et al.
Published: (2026)
Diffusion-Augmented Markov Decision Processes for Maximum Entropy Reinforcement Learning
by: Sanokowski, Sebastian, et al.
Published: (2025)
by: Sanokowski, Sebastian, et al.
Published: (2025)
Beyond Importance Sampling: Rejection-Gated Policy Optimization
by: Sun, Ziwu, et al.
Published: (2026)
by: Sun, Ziwu, et al.
Published: (2026)
Exploring Response Uncertainty in MLLMs: An Empirical Evaluation under Misleading Scenarios
by: Dang, Yunkai, et al.
Published: (2024)
by: Dang, Yunkai, et al.
Published: (2024)
Maximum Redundancy Pruning: A Principle-Driven Layerwise Sparsity Allocation for LLMs
by: Gao, Chang, et al.
Published: (2025)
by: Gao, Chang, et al.
Published: (2025)
Maximum Entropy Inverse Reinforcement Learning of Diffusion Models with Energy-Based Models
by: Yoon, Sangwoong, et al.
Published: (2024)
by: Yoon, Sangwoong, et al.
Published: (2024)
FLAC: Maximum Entropy RL via Kinetic Energy Regularized Bridge Matching
by: Lv, Lei, et al.
Published: (2026)
by: Lv, Lei, et al.
Published: (2026)
Boosting Maximum Entropy Reinforcement Learning via One-Step Flow Matching
by: Li, Zeqiao, et al.
Published: (2026)
by: Li, Zeqiao, et al.
Published: (2026)
EP-GRPO: Entropy-Progress Aligned Group Relative Policy Optimization with Implicit Process Guidance
by: Yu, Song, et al.
Published: (2026)
by: Yu, Song, et al.
Published: (2026)
Policy Gradient with Adaptive Entropy Annealing for Continual Fine-Tuning
by: Zhang, Yaqian, et al.
Published: (2026)
by: Zhang, Yaqian, et al.
Published: (2026)
Learning to Embed Distributions via Maximum Kernel Entropy
by: Kachaiev, Oleksii, et al.
Published: (2024)
by: Kachaiev, Oleksii, et al.
Published: (2024)
Learning Shortcuts: On the Misleading Promise of NLU in Language Models
by: Bihani, Geetanjali, et al.
Published: (2024)
by: Bihani, Geetanjali, et al.
Published: (2024)
Decision Flow Policy Optimization
by: Hu, Jifeng, et al.
Published: (2025)
by: Hu, Jifeng, et al.
Published: (2025)
Towards Flash Thinking via Decoupled Advantage Policy Optimization
by: Tan, Zezhong, et al.
Published: (2025)
by: Tan, Zezhong, et al.
Published: (2025)
Soft Adaptive Policy Optimization
by: Gao, Chang, et al.
Published: (2025)
by: Gao, Chang, et al.
Published: (2025)
Maximum Entropy Behavior Exploration for Sim2Real Zero-Shot Reinforcement Learning
by: Hu, Jiajun, et al.
Published: (2026)
by: Hu, Jiajun, et al.
Published: (2026)
Similar Items
-
Extremum-Seeking Action Selection for Accelerating Policy Optimization
by: Chang, Ya-Chien, et al.
Published: (2024) -
Improving Value Estimation Critically Enhances Vanilla Policy Gradient
by: Wang, Tao, et al.
Published: (2025) -
Learning Quadruped Walking from Seconds of Demonstration
by: Zhang, Ruipeng, et al.
Published: (2026) -
Fractal Landscapes in Policy Optimization
by: Wang, Tao, et al.
Published: (2023) -
Activation-Descent Regularization for Input Optimization of ReLU Networks
by: Yu, Hongzhan, et al.
Published: (2024)