Rethinking Entropy Interventions in RLVR: An Entropy Change Perspective
Fuente:
arXiv
Saved in:
| Main Authors: | Hao, Zhezheng, Wang, Hong, Liu, Haoyang, Luo, Jian, Yu, Jiarui, Dong, Hande, Lin, Qiang, Wang, Can, Chen, Jiawei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ReCreate: Reasoning and Creating Domain Agents Driven by Experience
by: Hao, Zhezheng, et al.
Published: (2026)
by: Hao, Zhezheng, et al.
Published: (2026)
LEPO: Latent Reasoning Policy Optimization for Large Language Models
by: Zhou, Yuyan, et al.
Published: (2026)
by: Zhou, Yuyan, et al.
Published: (2026)
Scheduling Your LLM Reinforcement Learning with Reasoning Trees
by: Wang, Hong, et al.
Published: (2025)
by: Wang, Hong, et al.
Published: (2025)
Evolve as a Team: Collaborative Self-Evolution for LLM-based Multi-Agent Systems
by: Hao, Zhezheng, et al.
Published: (2026)
by: Hao, Zhezheng, et al.
Published: (2026)
Exploration vs Exploitation: Rethinking RLVR through Clipping, Entropy, and Spurious Reward
by: Chen, Peter, et al.
Published: (2025)
by: Chen, Peter, et al.
Published: (2025)
Flexible Entropy Control in RLVR with a Gradient-Preserving Perspective
by: Chen, Kun, et al.
Published: (2026)
by: Chen, Kun, et al.
Published: (2026)
GAPO: Robust Advantage Estimation for Real-World Code LLMs
by: Zhang, Jianqing, et al.
Published: (2025)
by: Zhang, Jianqing, et al.
Published: (2025)
Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization
by: Xu, Huimin, et al.
Published: (2026)
by: Xu, Huimin, et al.
Published: (2026)
HEALing Entropy Collapse: Enhancing Exploration in Few-Shot RLVR via Hybrid-Domain Entropy Dynamics Alignment
by: Liu, Zhanyu, et al.
Published: (2026)
by: Liu, Zhanyu, et al.
Published: (2026)
Asymmetric Advantage Modulation Calibrates Entropy Dynamics in RLVR
by: Gu, Hengrui, et al.
Published: (2026)
by: Gu, Hengrui, et al.
Published: (2026)
Accelerating IC Thermal Simulation Data Generation via Block Krylov and Operator Action
by: Wang, Hong, et al.
Published: (2025)
by: Wang, Hong, et al.
Published: (2025)
SPREG: Structured Plan Repair with Entropy-Guided Test-Time Intervention for Large Language Model Reasoning
by: Wang, Xuan, et al.
Published: (2026)
by: Wang, Xuan, et al.
Published: (2026)
VL Norm: Rethink Loss Aggregation in RLVR
by: He, Zhiyuan, et al.
Published: (2025)
by: He, Zhiyuan, et al.
Published: (2025)
Rethinking Independent Cross-Entropy Loss For Graph-Structured Data
by: Miao, Rui, et al.
Published: (2024)
by: Miao, Rui, et al.
Published: (2024)
Rethinking Entropy Regularization in Large Reasoning Models
by: Jiang, Yuxian, et al.
Published: (2025)
by: Jiang, Yuxian, et al.
Published: (2025)
UniSVG: A Unified Dataset for Vector Graphic Understanding and Generation with Multimodal Large Language Models
by: Li, Jinke, et al.
Published: (2025)
by: Li, Jinke, et al.
Published: (2025)
SWE-Fuse: Empowering Software Agents via Issue-free Trajectory Learning and Entropy-aware RLVR Training
by: Wen, Xin-Cheng, et al.
Published: (2026)
by: Wen, Xin-Cheng, et al.
Published: (2026)
DSO: Dual-Scale Neural Operators for Stable Long-term Fluid Dynamics Forecasting
by: Dong, Huanshuo, et al.
Published: (2026)
by: Dong, Huanshuo, et al.
Published: (2026)
Rethinking GSPO: The Perplexity-Entropy Equivalence
by: Liu, Chi
Published: (2025)
by: Liu, Chi
Published: (2025)
Spurious Rewards: Rethinking Training Signals in RLVR
by: Shao, Rulin, et al.
Published: (2025)
by: Shao, Rulin, et al.
Published: (2025)
Rethinking Label Consistency of In-Context Learning: An Implicit Transductive Label Propagation Perspective
by: Chen, Haoyang, et al.
Published: (2025)
by: Chen, Haoyang, et al.
Published: (2025)
Structural Entropy Guided Probabilistic Coding
by: Huang, Xiang, et al.
Published: (2024)
by: Huang, Xiang, et al.
Published: (2024)
Rethinking Langevin Thompson Sampling from A Stochastic Approximation Perspective
by: Wang, Weixin, et al.
Published: (2025)
by: Wang, Weixin, et al.
Published: (2025)
Rethinking Entropy Minimization in Test-Time Adaptation for Autoregressive Models
by: Huang, Wei-Ping, et al.
Published: (2026)
by: Huang, Wei-Ping, et al.
Published: (2026)
The Unreasonable Effectiveness of Entropy Minimization in LLM Reasoning
by: Agarwal, Shivam, et al.
Published: (2025)
by: Agarwal, Shivam, et al.
Published: (2025)
PSL: Rethinking and Improving Softmax Loss from Pairwise Perspective for Recommendation
by: Yang, Weiqin, et al.
Published: (2024)
by: Yang, Weiqin, et al.
Published: (2024)
EntropyStop: Unsupervised Deep Outlier Detection with Loss Entropy
by: Huang, Yihong, et al.
Published: (2024)
by: Huang, Yihong, et al.
Published: (2024)
Mind Your Entropy: From Maximum Entropy to Trajectory Entropy-Constrained RL
by: Zhan, Guojian, et al.
Published: (2025)
by: Zhan, Guojian, et al.
Published: (2025)
Is Meta-Learning Out? Rethinking Unsupervised Few-Shot Classification with Limited Entropy
by: Guan, Yunchuan, et al.
Published: (2025)
by: Guan, Yunchuan, et al.
Published: (2025)
Maximum Entropy Reinforcement Learning with Diffusion Policy
by: Dong, Xiaoyi, et al.
Published: (2025)
by: Dong, Xiaoyi, et al.
Published: (2025)
Rethinking the Role of Entropy in Optimizing Tool-Use Behaviors for Large Language Model Agents
by: Li, Zeping, et al.
Published: (2026)
by: Li, Zeping, et al.
Published: (2026)
Enter the Void - Planning to Seek Entropy When Reward is Scarce
by: Sundar, Ashish, et al.
Published: (2025)
by: Sundar, Ashish, et al.
Published: (2025)
MEML-GRPO: Heterogeneous Multi-Expert Mutual Learning for RLVR Advancement
by: Jia, Weitao, et al.
Published: (2025)
by: Jia, Weitao, et al.
Published: (2025)
Intervention-Aware Forecasting: Breaking Historical Limits from a System Perspective
by: Xu, Zhijian, et al.
Published: (2024)
by: Xu, Zhijian, et al.
Published: (2024)
NoisyNN: Exploring the Impact of Information Entropy Change in Learning Systems
by: Yu, Xiaowei, et al.
Published: (2023)
by: Yu, Xiaowei, et al.
Published: (2023)
Unleashing the True Potential of LLMs: A Feedback-Triggered Self-Correction with Long-Term Multipath Decoding
by: Li, Jipeng, et al.
Published: (2025)
by: Li, Jipeng, et al.
Published: (2025)
Merino: Entropy-driven Design for Generative Language Models on IoT Devices
by: Zhao, Youpeng, et al.
Published: (2024)
by: Zhao, Youpeng, et al.
Published: (2024)
Entropy-Tree: Tree-Based Decoding with Entropy-Guided Exploration
by: Wei, Longxuan, et al.
Published: (2026)
by: Wei, Longxuan, et al.
Published: (2026)
Improving Entropy-Based Test-Time Adaptation from a Clustering View
by: Lin, Guoliang, et al.
Published: (2023)
by: Lin, Guoliang, et al.
Published: (2023)
Probing RLVR training instability through the lens of objective-level hacking
by: Dong, Yiming, et al.
Published: (2026)
by: Dong, Yiming, et al.
Published: (2026)
Similar Items
-
ReCreate: Reasoning and Creating Domain Agents Driven by Experience
by: Hao, Zhezheng, et al.
Published: (2026) -
LEPO: Latent Reasoning Policy Optimization for Large Language Models
by: Zhou, Yuyan, et al.
Published: (2026) -
Scheduling Your LLM Reinforcement Learning with Reasoning Trees
by: Wang, Hong, et al.
Published: (2025) -
Evolve as a Team: Collaborative Self-Evolution for LLM-based Multi-Agent Systems
by: Hao, Zhezheng, et al.
Published: (2026) -
Exploration vs Exploitation: Rethinking RLVR through Clipping, Entropy, and Spurious Reward
by: Chen, Peter, et al.
Published: (2025)