Mind Your Entropy: From Maximum Entropy to Trajectory Entropy-Constrained RL
Fuente:
arXiv
Saved in:
| Main Authors: | Zhan, Guojian, Wang, Likun, Wang, Pengcheng, Zhang, Feihong, Duan, Jingliang, Tomizuka, Masayoshi, Li, Shengbo Eben |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Bootstrap Off-policy with World Model
by: Zhan, Guojian, et al.
Published: (2025)
by: Zhan, Guojian, et al.
Published: (2025)
Distributional Soft Actor-Critic with Harmonic Gradient for Safe and Efficient Autonomous Driving in Multi-lane Scenarios
by: Zhang, Feihong, et al.
Published: (2025)
by: Zhang, Feihong, et al.
Published: (2025)
Diffusion Actor-Critic with Entropy Regulator
by: Wang, Yinuo, et al.
Published: (2024)
by: Wang, Yinuo, et al.
Published: (2024)
Mean Flow Policy with Instantaneous Velocity Constraint for One-step Action Generation
by: Zhan, Guojian, et al.
Published: (2026)
by: Zhan, Guojian, et al.
Published: (2026)
Enhanced DACER Algorithm with High Diffusion Efficiency
by: Wang, Yinuo, et al.
Published: (2025)
by: Wang, Yinuo, et al.
Published: (2025)
FastDSAC: Unlocking the Potential of Maximum Entropy RL in High-Dimensional Humanoid Control
by: Xue, Jun, et al.
Published: (2026)
by: Xue, Jun, et al.
Published: (2026)
Off-policy Reinforcement Learning with Model-based Exploration Augmentation
by: Wang, Likun, et al.
Published: (2025)
by: Wang, Likun, et al.
Published: (2025)
Distributional Soft Actor-Critic with Diffusion Policy
by: Liu, Tong, et al.
Published: (2025)
by: Liu, Tong, et al.
Published: (2025)
Maximum Entropy On-Policy Actor-Critic via Entropy Advantage Estimation
by: Choe, Jean Seong Bjorn, et al.
Published: (2024)
by: Choe, Jean Seong Bjorn, et al.
Published: (2024)
FLAC: Maximum Entropy RL via Kinetic Energy Regularized Bridge Matching
by: Lv, Lei, et al.
Published: (2026)
by: Lv, Lei, et al.
Published: (2026)
Reward-Punishment Reinforcement Learning with Maximum Entropy
by: Wang, Jiexin, et al.
Published: (2024)
by: Wang, Jiexin, et al.
Published: (2024)
When Maximum Entropy Misleads Policy Optimization
by: Zhang, Ruipeng, et al.
Published: (2025)
by: Zhang, Ruipeng, et al.
Published: (2025)
Maximum Entropy Reinforcement Learning with Diffusion Policy
by: Dong, Xiaoyi, et al.
Published: (2025)
by: Dong, Xiaoyi, et al.
Published: (2025)
On Entropy Control in LLM-RL Algorithms
by: Shen, Han
Published: (2025)
by: Shen, Han
Published: (2025)
Maximum Entropy Exploration Without the Rollouts
by: Adamczyk, Jacob, et al.
Published: (2026)
by: Adamczyk, Jacob, et al.
Published: (2026)
EntropyStop: Unsupervised Deep Outlier Detection with Loss Entropy
by: Huang, Yihong, et al.
Published: (2024)
by: Huang, Yihong, et al.
Published: (2024)
Rethinking Entropy Interventions in RLVR: An Entropy Change Perspective
by: Hao, Zhezheng, et al.
Published: (2025)
by: Hao, Zhezheng, et al.
Published: (2025)
Boosting Maximum Entropy Reinforcement Learning via One-Step Flow Matching
by: Li, Zeqiao, et al.
Published: (2026)
by: Li, Zeqiao, et al.
Published: (2026)
Curiosity & Entropy Driven Unsupervised RL in Multiple Environments
by: Dewan, Shaurya, et al.
Published: (2024)
by: Dewan, Shaurya, et al.
Published: (2024)
ELEMENT: Episodic and Lifelong Exploration via Maximum Entropy
by: Li, Hongming, et al.
Published: (2024)
by: Li, Hongming, et al.
Published: (2024)
Conformal Symplectic Optimization for Stable Reinforcement Learning
by: Lyu, Yao, et al.
Published: (2024)
by: Lyu, Yao, et al.
Published: (2024)
Deriving the Scaled-Dot-Function via Maximum Likelihood Estimation and Maximum Entropy Approach
by: Ma, Jiyong
Published: (2025)
by: Ma, Jiyong
Published: (2025)
Diffusion-Augmented Markov Decision Processes for Maximum Entropy Reinforcement Learning
by: Sanokowski, Sebastian, et al.
Published: (2025)
by: Sanokowski, Sebastian, et al.
Published: (2025)
Learning to Embed Distributions via Maximum Kernel Entropy
by: Kachaiev, Oleksii, et al.
Published: (2024)
by: Kachaiev, Oleksii, et al.
Published: (2024)
LaDi-RL: Latent Diffusion Reasoning Prevents Entropy Collapse in Reinforcement Learning
by: Kang, Haoqiang, et al.
Published: (2026)
by: Kang, Haoqiang, et al.
Published: (2026)
Predictive Lagrangian Optimization for Constrained Reinforcement Learning
by: Zhang, Tianqi, et al.
Published: (2025)
by: Zhang, Tianqi, et al.
Published: (2025)
EARL: Entropy-Aware RL Alignment of LLMs for Reliable RTL Code Generation
by: Shi, Jiahe, et al.
Published: (2025)
by: Shi, Jiahe, et al.
Published: (2025)
Maximum Entropy Inverse Reinforcement Learning of Diffusion Models with Energy-Based Models
by: Yoon, Sangwoong, et al.
Published: (2024)
by: Yoon, Sangwoong, et al.
Published: (2024)
Zeroth-Order Actor-Critic: An Evolutionary Framework for Sequential Decision Problems
by: Lei, Yuheng, et al.
Published: (2022)
by: Lei, Yuheng, et al.
Published: (2022)
Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization
by: Xu, Huimin, et al.
Published: (2026)
by: Xu, Huimin, et al.
Published: (2026)
TABES: Trajectory-Aware Backward-on-Entropy Steering for Masked Diffusion Models
by: Saini, Shreshth, et al.
Published: (2026)
by: Saini, Shreshth, et al.
Published: (2026)
Canonical Form of Datatic Description in Control Systems
by: Zhan, Guojian, et al.
Published: (2024)
by: Zhan, Guojian, et al.
Published: (2024)
Entropy-guided sequence weighting for efficient exploration in RL-based LLM fine-tuning
by: Vanlioglu, Abdullah
Published: (2025)
by: Vanlioglu, Abdullah
Published: (2025)
Maximum Entropy Behavior Exploration for Sim2Real Zero-Shot Reinforcement Learning
by: Hu, Jiajun, et al.
Published: (2026)
by: Hu, Jiajun, et al.
Published: (2026)
HEALing Entropy Collapse: Enhancing Exploration in Few-Shot RLVR via Hybrid-Domain Entropy Dynamics Alignment
by: Liu, Zhanyu, et al.
Published: (2026)
by: Liu, Zhanyu, et al.
Published: (2026)
Hard Prompts Made Interpretable: Sparse Entropy Regularization for Prompt Tuning with RL
by: Choi, Yunseon, et al.
Published: (2024)
by: Choi, Yunseon, et al.
Published: (2024)
HEAL: Hindsight Entropy-Assisted Learning for Reasoning Distillation
by: Zhang, Wenjing, et al.
Published: (2026)
by: Zhang, Wenjing, et al.
Published: (2026)
$α$-GAN by Rényi Cross Entropy
by: Ding, Ni, et al.
Published: (2025)
by: Ding, Ni, et al.
Published: (2025)
What Scales in Cross-Entropy Scaling Law?
by: Yan, Junxi, et al.
Published: (2025)
by: Yan, Junxi, et al.
Published: (2025)
Entropy-Preserving Reinforcement Learning
by: Petrenko, Aleksei, et al.
Published: (2026)
by: Petrenko, Aleksei, et al.
Published: (2026)
Similar Items
-
Bootstrap Off-policy with World Model
by: Zhan, Guojian, et al.
Published: (2025) -
Distributional Soft Actor-Critic with Harmonic Gradient for Safe and Efficient Autonomous Driving in Multi-lane Scenarios
by: Zhang, Feihong, et al.
Published: (2025) -
Diffusion Actor-Critic with Entropy Regulator
by: Wang, Yinuo, et al.
Published: (2024) -
Mean Flow Policy with Instantaneous Velocity Constraint for One-step Action Generation
by: Zhan, Guojian, et al.
Published: (2026) -
Enhanced DACER Algorithm with High Diffusion Efficiency
by: Wang, Yinuo, et al.
Published: (2025)