Beyond High-Entropy Exploration: Correctness-Aware Low-Entropy Segment-Based Advantage Shaping for Reasoning LLMs
Fuente:
arXiv
Salvato in:
| Autori principali: | Chen, Xinzhu, Li, Xuesheng, Sun, Zhongxiang, Yu, Weijie |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Maximum Entropy On-Policy Actor-Critic via Entropy Advantage Estimation
di: Choe, Jean Seong Bjorn, et al.
Pubblicazione: (2024)
di: Choe, Jean Seong Bjorn, et al.
Pubblicazione: (2024)
ELEMENT: Episodic and Lifelong Exploration via Maximum Entropy
di: Li, Hongming, et al.
Pubblicazione: (2024)
di: Li, Hongming, et al.
Pubblicazione: (2024)
Entropy-Aware Model Initialization for Effective Exploration in Deep Reinforcement Learning
di: Jang, Sooyoung, et al.
Pubblicazione: (2021)
di: Jang, Sooyoung, et al.
Pubblicazione: (2021)
Reinforcement Learning for Diffusion LLMs with Entropy-Guided Step Selection and Stepwise Advantages
di: Kunde, Vishnu Teja, et al.
Pubblicazione: (2026)
di: Kunde, Vishnu Teja, et al.
Pubblicazione: (2026)
Maximum Entropy Exploration Without the Rollouts
di: Adamczyk, Jacob, et al.
Pubblicazione: (2026)
di: Adamczyk, Jacob, et al.
Pubblicazione: (2026)
Entropy-Guided Loop: Achieving Reasoning through Uncertainty-Aware Generation
di: Correa, Andrew G. A., et al.
Pubblicazione: (2025)
di: Correa, Andrew G. A., et al.
Pubblicazione: (2025)
Asymmetric Advantage Modulation Calibrates Entropy Dynamics in RLVR
di: Gu, Hengrui, et al.
Pubblicazione: (2026)
di: Gu, Hengrui, et al.
Pubblicazione: (2026)
HEALing Entropy Collapse: Enhancing Exploration in Few-Shot RLVR via Hybrid-Domain Entropy Dynamics Alignment
di: Liu, Zhanyu, et al.
Pubblicazione: (2026)
di: Liu, Zhanyu, et al.
Pubblicazione: (2026)
Rethinking Entropy Regularization in Large Reasoning Models
di: Jiang, Yuxian, et al.
Pubblicazione: (2025)
di: Jiang, Yuxian, et al.
Pubblicazione: (2025)
EARL: Entropy-Aware RL Alignment of LLMs for Reliable RTL Code Generation
di: Shi, Jiahe, et al.
Pubblicazione: (2025)
di: Shi, Jiahe, et al.
Pubblicazione: (2025)
EntroPIC: Towards Stable Long-Term Training of LLMs via Entropy Stabilization with Proportional-Integral Control
di: Yang, Kai, et al.
Pubblicazione: (2025)
di: Yang, Kai, et al.
Pubblicazione: (2025)
Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning
di: Wang, Shenzhi, et al.
Pubblicazione: (2025)
di: Wang, Shenzhi, et al.
Pubblicazione: (2025)
ERPO: Token-Level Entropy-Regulated Policy Optimization for Large Reasoning Models
di: Yu, Song, et al.
Pubblicazione: (2026)
di: Yu, Song, et al.
Pubblicazione: (2026)
Rethinking Entropy Interventions in RLVR: An Entropy Change Perspective
di: Hao, Zhezheng, et al.
Pubblicazione: (2025)
di: Hao, Zhezheng, et al.
Pubblicazione: (2025)
No Prompt Left Behind: Exploiting Zero-Variance Prompts in LLM Reinforcement Learning via Entropy-Guided Advantage Shaping
di: Le, Thanh-Long V., et al.
Pubblicazione: (2025)
di: Le, Thanh-Long V., et al.
Pubblicazione: (2025)
Accelerating Reinforcement Learning with Value-Conditional State Entropy Exploration
di: Kim, Dongyoung, et al.
Pubblicazione: (2023)
di: Kim, Dongyoung, et al.
Pubblicazione: (2023)
Mind Your Entropy: From Maximum Entropy to Trajectory Entropy-Constrained RL
di: Zhan, Guojian, et al.
Pubblicazione: (2025)
di: Zhan, Guojian, et al.
Pubblicazione: (2025)
Entropy-Guided Data-Efficient Training for Multimodal Reasoning Reward Models
di: Yang, Shidong, et al.
Pubblicazione: (2026)
di: Yang, Shidong, et al.
Pubblicazione: (2026)
HEAL: Hindsight Entropy-Assisted Learning for Reasoning Distillation
di: Zhang, Wenjing, et al.
Pubblicazione: (2026)
di: Zhang, Wenjing, et al.
Pubblicazione: (2026)
Exploration vs Exploitation: Rethinking RLVR through Clipping, Entropy, and Spurious Reward
di: Chen, Peter, et al.
Pubblicazione: (2025)
di: Chen, Peter, et al.
Pubblicazione: (2025)
The Unreasonable Effectiveness of Entropy Minimization in LLM Reasoning
di: Agarwal, Shivam, et al.
Pubblicazione: (2025)
di: Agarwal, Shivam, et al.
Pubblicazione: (2025)
Adaptive Ensembles of Fine-Tuned Transformers for LLM-Generated Text Detection
di: Lai, Zhixin, et al.
Pubblicazione: (2024)
di: Lai, Zhixin, et al.
Pubblicazione: (2024)
The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models
di: Cui, Ganqu, et al.
Pubblicazione: (2025)
di: Cui, Ganqu, et al.
Pubblicazione: (2025)
A Multi-dimensional Semantic Surprise Framework Based on Low-Entropy Semantic Manifolds for Fine-Grained Out-of-Distribution Detection
di: Peng, Ningkang, et al.
Pubblicazione: (2025)
di: Peng, Ningkang, et al.
Pubblicazione: (2025)
Stable Adaptive Thinking via Advantage Shaping and Length-Aware Gradient Regulation
di: Xu, Zihang, et al.
Pubblicazione: (2026)
di: Xu, Zihang, et al.
Pubblicazione: (2026)
D-PACE: Dynamic Position-Aware Cross-Entropy for Parallel Speculative Drafting
di: Wu, Tianyu, et al.
Pubblicazione: (2026)
di: Wu, Tianyu, et al.
Pubblicazione: (2026)
Unveiling Implicit Advantage Symmetry: Why GRPO Struggles with Exploration and Difficulty Adaptation
di: Yu, Zhiqi, et al.
Pubblicazione: (2026)
di: Yu, Zhiqi, et al.
Pubblicazione: (2026)
Memory-Based Advantage Shaping for LLM-Guided Reinforcement Learning
di: Nourzad, Narjes, et al.
Pubblicazione: (2026)
di: Nourzad, Narjes, et al.
Pubblicazione: (2026)
Improving Offline-to-Online Reinforcement Learning with Q Conditioned State Entropy Exploration
di: Zhang, Ziqi, et al.
Pubblicazione: (2023)
di: Zhang, Ziqi, et al.
Pubblicazione: (2023)
LLMs Uncertainty Quantification via Adaptive Conformal Semantic Entropy
di: Karimi, Hamed, et al.
Pubblicazione: (2026)
di: Karimi, Hamed, et al.
Pubblicazione: (2026)
Tracking Drift: Variation-Aware Entropy Scheduling for Non-Stationary Reinforcement Learning
di: Wang, Tongxi, et al.
Pubblicazione: (2026)
di: Wang, Tongxi, et al.
Pubblicazione: (2026)
Accelerating RL for LLM Reasoning with Optimal Advantage Regression
di: Brantley, Kianté, et al.
Pubblicazione: (2025)
di: Brantley, Kianté, et al.
Pubblicazione: (2025)
From Implicit Exploration to Structured Reasoning: Leveraging Guideline and Refinement for LLMs
di: Chen, Jiaxiang, et al.
Pubblicazione: (2025)
di: Chen, Jiaxiang, et al.
Pubblicazione: (2025)
Beyond Correctness: Confidence-Aware Reward Modeling for Enhancing Large Language Model Reasoning
di: He, Qianxi, et al.
Pubblicazione: (2025)
di: He, Qianxi, et al.
Pubblicazione: (2025)
FLAC: Maximum Entropy RL via Kinetic Energy Regularized Bridge Matching
di: Lv, Lei, et al.
Pubblicazione: (2026)
di: Lv, Lei, et al.
Pubblicazione: (2026)
Maximum Entropy Behavior Exploration for Sim2Real Zero-Shot Reinforcement Learning
di: Hu, Jiajun, et al.
Pubblicazione: (2026)
di: Hu, Jiajun, et al.
Pubblicazione: (2026)
Revisiting Transformers through the Lens of Low Entropy and Dynamic Sparsity
di: Ren, Ruifeng, et al.
Pubblicazione: (2025)
di: Ren, Ruifeng, et al.
Pubblicazione: (2025)
ACE : Off-Policy Actor-Critic with Causality-Aware Entropy Regularization
di: Ji, Tianying, et al.
Pubblicazione: (2024)
di: Ji, Tianying, et al.
Pubblicazione: (2024)
On the Entropy Dynamics in Reinforcement Fine-Tuning of Large Language Models
di: Wang, Shumin, et al.
Pubblicazione: (2026)
di: Wang, Shumin, et al.
Pubblicazione: (2026)
EntropyStop: Unsupervised Deep Outlier Detection with Loss Entropy
di: Huang, Yihong, et al.
Pubblicazione: (2024)
di: Huang, Yihong, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Maximum Entropy On-Policy Actor-Critic via Entropy Advantage Estimation
di: Choe, Jean Seong Bjorn, et al.
Pubblicazione: (2024) -
ELEMENT: Episodic and Lifelong Exploration via Maximum Entropy
di: Li, Hongming, et al.
Pubblicazione: (2024) -
Entropy-Aware Model Initialization for Effective Exploration in Deep Reinforcement Learning
di: Jang, Sooyoung, et al.
Pubblicazione: (2021) -
Reinforcement Learning for Diffusion LLMs with Entropy-Guided Step Selection and Stepwise Advantages
di: Kunde, Vishnu Teja, et al.
Pubblicazione: (2026) -
Maximum Entropy Exploration Without the Rollouts
di: Adamczyk, Jacob, et al.
Pubblicazione: (2026)