FastDSAC: Unlocking the Potential of Maximum Entropy RL in High-Dimensional Humanoid Control
Fuente:
arXiv
Saved in:
| Main Authors: | Xue, Jun, Wang, Junze, Wang, Shanze, Zhang, Xinming, Chen, Yanjun, Zhang, Wei |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Rethinking Soft Actor-Critic in High-Dimensional Action Spaces: The Cost of Ignoring Distribution Shift
by: Chen, Yanjun, et al.
Published: (2024)
by: Chen, Yanjun, et al.
Published: (2024)
Mind Your Entropy: From Maximum Entropy to Trajectory Entropy-Constrained RL
by: Zhan, Guojian, et al.
Published: (2025)
by: Zhan, Guojian, et al.
Published: (2025)
DSAC: Distributional Soft Actor-Critic for Risk-Sensitive Reinforcement Learning
by: Ma, Xiaoteng, et al.
Published: (2020)
by: Ma, Xiaoteng, et al.
Published: (2020)
Exact Is Easier: Credit Assignment for Cooperative LLM Agents
by: Chen, Yanjun, et al.
Published: (2026)
by: Chen, Yanjun, et al.
Published: (2026)
Revisiting Entropy Regularization: Adaptive Coefficient Unlocks Its Potential for LLM Reinforcement Learning
by: Zhang, Xiaoyun, et al.
Published: (2025)
by: Zhang, Xiaoyun, et al.
Published: (2025)
FLAC: Maximum Entropy RL via Kinetic Energy Regularized Bridge Matching
by: Lv, Lei, et al.
Published: (2026)
by: Lv, Lei, et al.
Published: (2026)
On Entropy Control in LLM-RL Algorithms
by: Shen, Han
Published: (2025)
by: Shen, Han
Published: (2025)
Angles Don't Lie: Unlocking Training-Efficient RL Through the Model's Own Signals
by: Wang, Qinsi, et al.
Published: (2025)
by: Wang, Qinsi, et al.
Published: (2025)
When Maximum Entropy Misleads Policy Optimization
by: Zhang, Ruipeng, et al.
Published: (2025)
by: Zhang, Ruipeng, et al.
Published: (2025)
Maximum Entropy Reinforcement Learning with Diffusion Policy
by: Dong, Xiaoyi, et al.
Published: (2025)
by: Dong, Xiaoyi, et al.
Published: (2025)
DSAC-C: Constrained Maximum Entropy for Robust Discrete Soft-Actor Critic
by: Neo, Dexter, et al.
Published: (2023)
by: Neo, Dexter, et al.
Published: (2023)
FastTD3: Simple, Fast, and Capable Reinforcement Learning for Humanoid Control
by: Seo, Younggyo, et al.
Published: (2025)
by: Seo, Younggyo, et al.
Published: (2025)
Reward-Punishment Reinforcement Learning with Maximum Entropy
by: Wang, Jiexin, et al.
Published: (2024)
by: Wang, Jiexin, et al.
Published: (2024)
Embodiment-Aware Generalist Specialist Distillation for Unified Humanoid Whole-Body Control
by: Peng, Quanquan, et al.
Published: (2026)
by: Peng, Quanquan, et al.
Published: (2026)
TDMPBC: Self-Imitative Reinforcement Learning for Humanoid Robot Control
by: Zhuang, Zifeng, et al.
Published: (2025)
by: Zhuang, Zifeng, et al.
Published: (2025)
Boosting Maximum Entropy Reinforcement Learning via One-Step Flow Matching
by: Li, Zeqiao, et al.
Published: (2026)
by: Li, Zeqiao, et al.
Published: (2026)
Mixtures of Experts Unlock Parameter Scaling for Deep RL
by: Obando-Ceron, Johan, et al.
Published: (2024)
by: Obando-Ceron, Johan, et al.
Published: (2024)
Task-Aware Adaptive Modulation: A Replay-Free and Resource-Efficient Approach For Continual Graph Learning
by: Liu, Jingtao, et al.
Published: (2025)
by: Liu, Jingtao, et al.
Published: (2025)
Maximum Entropy Exploration Without the Rollouts
by: Adamczyk, Jacob, et al.
Published: (2026)
by: Adamczyk, Jacob, et al.
Published: (2026)
Knapsack RL: Unlocking Exploration of LLMs via Optimizing Budget Allocation
by: Li, Ziniu, et al.
Published: (2025)
by: Li, Ziniu, et al.
Published: (2025)
ArenaRL: Scaling RL for Open-Ended Agents via Tournament-based Relative Ranking
by: Zhang, Qiang, et al.
Published: (2026)
by: Zhang, Qiang, et al.
Published: (2026)
Learning Humanoid Standing-up Control across Diverse Postures
by: Huang, Tao, et al.
Published: (2025)
by: Huang, Tao, et al.
Published: (2025)
Maximum Entropy On-Policy Actor-Critic via Entropy Advantage Estimation
by: Choe, Jean Seong Bjorn, et al.
Published: (2024)
by: Choe, Jean Seong Bjorn, et al.
Published: (2024)
Beyond Token-level Supervision: Unlocking the Potential of Decoding-based Regression via Reinforcement Learning
by: Chen, Ming, et al.
Published: (2025)
by: Chen, Ming, et al.
Published: (2025)
QaRL: Rollout-Aligned Quantization-Aware RL for Fast and Stable Training under Training--Inference Mismatch
by: Gu, Hao, et al.
Published: (2026)
by: Gu, Hao, et al.
Published: (2026)
Unlocking Biomedical Insights: Hierarchical Attention Networks for High-Dimensional Data Interpretation
by: Nair, Rekha R, et al.
Published: (2025)
by: Nair, Rekha R, et al.
Published: (2025)
Deriving the Scaled-Dot-Function via Maximum Likelihood Estimation and Maximum Entropy Approach
by: Ma, Jiyong
Published: (2025)
by: Ma, Jiyong
Published: (2025)
LaDi-RL: Latent Diffusion Reasoning Prevents Entropy Collapse in Reinforcement Learning
by: Kang, Haoqiang, et al.
Published: (2026)
by: Kang, Haoqiang, et al.
Published: (2026)
MuJoCo MPC for Humanoid Control: Evaluation on HumanoidBench
by: Meser, Moritz, et al.
Published: (2024)
by: Meser, Moritz, et al.
Published: (2024)
Network Sparsity Unlocks the Scaling Potential of Deep Reinforcement Learning
by: Ma, Guozheng, et al.
Published: (2025)
by: Ma, Guozheng, et al.
Published: (2025)
Curiosity & Entropy Driven Unsupervised RL in Multiple Environments
by: Dewan, Shaurya, et al.
Published: (2024)
by: Dewan, Shaurya, et al.
Published: (2024)
ELEMENT: Episodic and Lifelong Exploration via Maximum Entropy
by: Li, Hongming, et al.
Published: (2024)
by: Li, Hongming, et al.
Published: (2024)
Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training
by: Liu, Mingjie, et al.
Published: (2025)
by: Liu, Mingjie, et al.
Published: (2025)
VAM: Verbalized Action Masking for Controllable Exploration in RL Post-Training -- A Chess Case Study
by: Zhang, Zhicheng, et al.
Published: (2026)
by: Zhang, Zhicheng, et al.
Published: (2026)
EntropyStop: Unsupervised Deep Outlier Detection with Loss Entropy
by: Huang, Yihong, et al.
Published: (2024)
by: Huang, Yihong, et al.
Published: (2024)
Humanoid-inspired Causal Representation Learning for Domain Generalization
by: Tao, Ze, et al.
Published: (2025)
by: Tao, Ze, et al.
Published: (2025)
FASTER: Value-Guided Sampling for Fast RL
by: Dong, Perry, et al.
Published: (2026)
by: Dong, Perry, et al.
Published: (2026)
GAS: Enhancing Reward-Cost Balance of Generative Model-assisted Offline Safe RL
by: Liu, Zifan, et al.
Published: (2026)
by: Liu, Zifan, et al.
Published: (2026)
MiMo: Unlocking the Reasoning Potential of Language Model -- From Pretraining to Posttraining
by: Xiaomi, LLM-Core, et al.
Published: (2025)
by: Xiaomi, LLM-Core, et al.
Published: (2025)
The Optimal Token Baseline: Variance Reduction for Long-Horizon LLM-RL
by: Li, Yingru, et al.
Published: (2026)
by: Li, Yingru, et al.
Published: (2026)
Similar Items
-
Rethinking Soft Actor-Critic in High-Dimensional Action Spaces: The Cost of Ignoring Distribution Shift
by: Chen, Yanjun, et al.
Published: (2024) -
Mind Your Entropy: From Maximum Entropy to Trajectory Entropy-Constrained RL
by: Zhan, Guojian, et al.
Published: (2025) -
DSAC: Distributional Soft Actor-Critic for Risk-Sensitive Reinforcement Learning
by: Ma, Xiaoteng, et al.
Published: (2020) -
Exact Is Easier: Credit Assignment for Cooperative LLM Agents
by: Chen, Yanjun, et al.
Published: (2026) -
Revisiting Entropy Regularization: Adaptive Coefficient Unlocks Its Potential for LLM Reinforcement Learning
by: Zhang, Xiaoyun, et al.
Published: (2025)