Addressing Performance Saturation for LLM RL via Precise Entropy Curve Control
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Bolian, Wang, Yifan, Ding, Yi, Lochab, Anamika, Grama, Ananth, Zhang, Ruqi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Cascade Reward Sampling for Efficient Decoding-Time Alignment
von: Li, Bolian, et al.
Veröffentlicht: (2024)
von: Li, Bolian, et al.
Veröffentlicht: (2024)
Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity
von: Lochab, Anamika, et al.
Veröffentlicht: (2026)
von: Lochab, Anamika, et al.
Veröffentlicht: (2026)
Energy-Based Reward Models for Robust Language Model Alignment
von: Lochab, Anamika, et al.
Veröffentlicht: (2025)
von: Lochab, Anamika, et al.
Veröffentlicht: (2025)
VERA-V: Variational Inference Framework for Jailbreaking Vision-Language Models
von: Liao, Qilin, et al.
Veröffentlicht: (2025)
von: Liao, Qilin, et al.
Veröffentlicht: (2025)
ETA: Evaluating Then Aligning Safety of Vision Language Models at Inference Time
von: Ding, Yi, et al.
Veröffentlicht: (2024)
von: Ding, Yi, et al.
Veröffentlicht: (2024)
VERA: Variational Inference Framework for Jailbreaking Large Language Models
von: Lochab, Anamika, et al.
Veröffentlicht: (2025)
von: Lochab, Anamika, et al.
Veröffentlicht: (2025)
Learning Self-Correction in Vision-Language Models via Rollout Augmentation
von: Ding, Yi, et al.
Veröffentlicht: (2026)
von: Ding, Yi, et al.
Veröffentlicht: (2026)
Entropy-MCMC: Sampling from Flat Basins with Ease
von: Li, Bolian, et al.
Veröffentlicht: (2023)
von: Li, Bolian, et al.
Veröffentlicht: (2023)
Making Reliable and Flexible Decisions in Long-tailed Classification
von: Li, Bolian, et al.
Veröffentlicht: (2025)
von: Li, Bolian, et al.
Veröffentlicht: (2025)
Controlled LLM Decoding via Discrete Auto-regressive Biasing
von: Pynadath, Patrick, et al.
Veröffentlicht: (2025)
von: Pynadath, Patrick, et al.
Veröffentlicht: (2025)
DRIFT: Learning from Abundant User Dissatisfaction in Real-World Preference Learning
von: Wang, Yifan, et al.
Veröffentlicht: (2025)
von: Wang, Yifan, et al.
Veröffentlicht: (2025)
Sherlock: Self-Correcting Reasoning in Vision-Language Models
von: Ding, Yi, et al.
Veröffentlicht: (2025)
von: Ding, Yi, et al.
Veröffentlicht: (2025)
SARL: Label-Free Reinforcement Learning by Rewarding Reasoning Topology
von: Wang, Yifan, et al.
Veröffentlicht: (2026)
von: Wang, Yifan, et al.
Veröffentlicht: (2026)
Inference-Time Code Selection via Symbolic Equivalence Partitioning
von: Cho, David, et al.
Veröffentlicht: (2026)
von: Cho, David, et al.
Veröffentlicht: (2026)
Bayesian Computation in Deep Learning
von: Chen, Wenlong, et al.
Veröffentlicht: (2025)
von: Chen, Wenlong, et al.
Veröffentlicht: (2025)
Deconvolving Complex Neuronal Networks into Interpretable Task-Specific Connectomes
von: Wang, Yifan, et al.
Veröffentlicht: (2024)
von: Wang, Yifan, et al.
Veröffentlicht: (2024)
CoT-UQ: Improving Response-wise Uncertainty Quantification in LLMs with Chain-of-Thought
von: Zhang, Boxuan, et al.
Veröffentlicht: (2025)
von: Zhang, Boxuan, et al.
Veröffentlicht: (2025)
Save the Good Prefix: Precise Error Penalization via Process-Supervised RL to Enhance LLM Reasoning
von: Liu, Haolin, et al.
Veröffentlicht: (2026)
von: Liu, Haolin, et al.
Veröffentlicht: (2026)
FP8-RL: A Practical and Stable Low-Precision Stack for LLM Reinforcement Learning
von: Qiu, Zhaopeng, et al.
Veröffentlicht: (2026)
von: Qiu, Zhaopeng, et al.
Veröffentlicht: (2026)
Why Any-Order Autoregressive Models Need Two-Stream Attention: A Structural-Semantic Tradeoff
von: Pynadath, Patrick, et al.
Veröffentlicht: (2026)
von: Pynadath, Patrick, et al.
Veröffentlicht: (2026)
SPA-RL: Reinforcing LLM Agents via Stepwise Progress Attribution
von: Wang, Hanlin, et al.
Veröffentlicht: (2025)
von: Wang, Hanlin, et al.
Veröffentlicht: (2025)
Entropy Law: The Story Behind Data Compression and LLM Performance
von: Yin, Mingjia, et al.
Veröffentlicht: (2024)
von: Yin, Mingjia, et al.
Veröffentlicht: (2024)
Stacey: Promoting Stochastic Steepest Descent via Accelerated $\ell_p$-Smooth Nonconvex Optimization
von: Luo, Xinyu, et al.
Veröffentlicht: (2025)
von: Luo, Xinyu, et al.
Veröffentlicht: (2025)
Robust Online Classification: From Estimation to Denoising
von: Wu, Changlong, et al.
Veröffentlicht: (2023)
von: Wu, Changlong, et al.
Veröffentlicht: (2023)
Two Calls, Two Moments, and the Vote-Accuracy Curve of Repeated LLM Inference
von: Liu, Yi
Veröffentlicht: (2026)
von: Liu, Yi
Veröffentlicht: (2026)
Generative Frontiers: Why Evaluation Matters for Diffusion Language Models
von: Pynadath, Patrick, et al.
Veröffentlicht: (2026)
von: Pynadath, Patrick, et al.
Veröffentlicht: (2026)
Beyond Markovian: Reflective Exploration via Bayes-Adaptive RL for LLM Reasoning
von: Zhang, Shenao, et al.
Veröffentlicht: (2025)
von: Zhang, Shenao, et al.
Veröffentlicht: (2025)
Memex(RL): Scaling Long-Horizon LLM Agents via Indexed Experience Memory
von: Wang, Zhenting, et al.
Veröffentlicht: (2026)
von: Wang, Zhenting, et al.
Veröffentlicht: (2026)
ResRL: Boosting LLM Reasoning via Negative Sample Projection Residual Reinforcement Learning
von: Lin, Zihan, et al.
Veröffentlicht: (2026)
von: Lin, Zihan, et al.
Veröffentlicht: (2026)
Explore Briefly, Then Decide: Mitigating LLM Overthinking via Cumulative Entropy Regulation
von: Bin, Yi, et al.
Veröffentlicht: (2025)
von: Bin, Yi, et al.
Veröffentlicht: (2025)
No Free Lunch: Fundamental Limits of Learning Non-Hallucinating Generative Models
von: Wu, Changlong, et al.
Veröffentlicht: (2024)
von: Wu, Changlong, et al.
Veröffentlicht: (2024)
Generalized Learning of Coefficients in Spectral Graph Convolutional Networks
von: Coşkun, Mustafa, et al.
Veröffentlicht: (2024)
von: Coşkun, Mustafa, et al.
Veröffentlicht: (2024)
DISA: Offline Importance Sampling for Distribution-Matching LLM-RL
von: Wang, Shaobo, et al.
Veröffentlicht: (2026)
von: Wang, Shaobo, et al.
Veröffentlicht: (2026)
DuoGuard: A Two-Player RL-Driven Framework for Multilingual LLM Guardrails
von: Deng, Yihe, et al.
Veröffentlicht: (2025)
von: Deng, Yihe, et al.
Veröffentlicht: (2025)
Reward-Shifted Speculative Sampling Is An Efficient Test-Time Weak-to-Strong Aligner
von: Li, Bolian, et al.
Veröffentlicht: (2025)
von: Li, Bolian, et al.
Veröffentlicht: (2025)
Chart-RL: Generalized Chart Comprehension via Reinforcement Learning with Verifiable Rewards
von: Zhang, Xin, et al.
Veröffentlicht: (2026)
von: Zhang, Xin, et al.
Veröffentlicht: (2026)
On Entropy Control in LLM-RL Algorithms
von: Shen, Han
Veröffentlicht: (2025)
von: Shen, Han
Veröffentlicht: (2025)
Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents
von: Wang, Jiawei, et al.
Veröffentlicht: (2025)
von: Wang, Jiawei, et al.
Veröffentlicht: (2025)
FlowRL: Matching Reward Distributions for LLM Reasoning
von: Zhu, Xuekai, et al.
Veröffentlicht: (2025)
von: Zhu, Xuekai, et al.
Veröffentlicht: (2025)
Sparse-RL: Breaking the Memory Wall in LLM Reinforcement Learning via Stable Sparse Rollouts
von: Luo, Sijia, et al.
Veröffentlicht: (2026)
von: Luo, Sijia, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Cascade Reward Sampling for Efficient Decoding-Time Alignment
von: Li, Bolian, et al.
Veröffentlicht: (2024) -
Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity
von: Lochab, Anamika, et al.
Veröffentlicht: (2026) -
Energy-Based Reward Models for Robust Language Model Alignment
von: Lochab, Anamika, et al.
Veröffentlicht: (2025) -
VERA-V: Variational Inference Framework for Jailbreaking Vision-Language Models
von: Liao, Qilin, et al.
Veröffentlicht: (2025) -
ETA: Evaluating Then Aligning Safety of Vision Language Models at Inference Time
von: Ding, Yi, et al.
Veröffentlicht: (2024)