Saved in:
| Main Authors: | Agarwal, Shivam, Zhang, Zimin, Yuan, Lifan, Han, Jiawei, Peng, Hao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2505.15134 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Unreasonable Effectiveness of Eccentric Automatic Prompts
by: Battle, Rick, et al.
Published: (2024)
by: Battle, Rick, et al.
Published: (2024)
Beyond Semantics: The Unreasonable Effectiveness of Reasonless Intermediate Tokens
by: Valmeekam, Karthik, et al.
Published: (2025)
by: Valmeekam, Karthik, et al.
Published: (2025)
The Unreasonable Effectiveness of Model Merging for Cross-Lingual Transfer in LLMs
by: Bandarkar, Lucas, et al.
Published: (2025)
by: Bandarkar, Lucas, et al.
Published: (2025)
The Unreasonable Effectiveness of Easy Training Data for Hard Tasks
by: Hase, Peter, et al.
Published: (2024)
by: Hase, Peter, et al.
Published: (2024)
The Unreasonable Effectiveness of Open Science in AI: A Replication Study
by: Gundersen, Odd Erik, et al.
Published: (2024)
by: Gundersen, Odd Erik, et al.
Published: (2024)
The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models
by: Cui, Ganqu, et al.
Published: (2025)
by: Cui, Ganqu, et al.
Published: (2025)
The Unreasonable Effectiveness of Discrete-Time Gaussian Process Mixtures for Robot Policy Learning
by: von Hartz, Jan Ole, et al.
Published: (2025)
by: von Hartz, Jan Ole, et al.
Published: (2025)
Do We Need Adam? Surprisingly Strong and Sparse Reinforcement Learning with SGD in LLMs
by: Mukherjee, Sagnik, et al.
Published: (2026)
by: Mukherjee, Sagnik, et al.
Published: (2026)
Curiosity & Entropy Driven Unsupervised RL in Multiple Environments
by: Dewan, Shaurya, et al.
Published: (2024)
by: Dewan, Shaurya, et al.
Published: (2024)
Advancing LLM Reasoning Generalists with Preference Trees
by: Yuan, Lifan, et al.
Published: (2024)
by: Yuan, Lifan, et al.
Published: (2024)
Examining LLMs' Uncertainty Expression Towards Questions Outside Parametric Knowledge
by: Liu, Genglin, et al.
Published: (2023)
by: Liu, Genglin, et al.
Published: (2023)
On Entropy Control in LLM-RL Algorithms
by: Shen, Han
Published: (2025)
by: Shen, Han
Published: (2025)
PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment
by: Li, Jiawei, et al.
Published: (2024)
by: Li, Jiawei, et al.
Published: (2024)
CRAFT: Customizing LLMs by Creating and Retrieving from Specialized Toolsets
by: Yuan, Lifan, et al.
Published: (2023)
by: Yuan, Lifan, et al.
Published: (2023)
Rethinking Channel Dependence for Multivariate Time Series Forecasting: Learning from Leading Indicators
by: Zhao, Lifan, et al.
Published: (2024)
by: Zhao, Lifan, et al.
Published: (2024)
Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning
by: Wang, Shenzhi, et al.
Published: (2025)
by: Wang, Shenzhi, et al.
Published: (2025)
Rethinking Entropy Interventions in RLVR: An Entropy Change Perspective
by: Hao, Zhezheng, et al.
Published: (2025)
by: Hao, Zhezheng, et al.
Published: (2025)
MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback
by: Wang, Xingyao, et al.
Published: (2023)
by: Wang, Xingyao, et al.
Published: (2023)
Trajectory Bellman Residual Minimization: A Simple Value-Based Method for LLM Reasoning
by: Yuan, Yurun, et al.
Published: (2025)
by: Yuan, Yurun, et al.
Published: (2025)
MPC-Minimized Secure LLM Inference
by: Rathee, Deevashwer, et al.
Published: (2024)
by: Rathee, Deevashwer, et al.
Published: (2024)
Stable Reasoning, Unstable Responses: Mitigating LLM Deception via Stability Asymmetry
by: Zhang, Guoxi, et al.
Published: (2026)
by: Zhang, Guoxi, et al.
Published: (2026)
Bridging Internal Probability and Self-Consistency for Effective and Efficient LLM Reasoning
by: Zhou, Zhi, et al.
Published: (2025)
by: Zhou, Zhi, et al.
Published: (2025)
UDora: A Unified Red Teaming Framework against LLM Agents by Dynamically Hijacking Their Own Reasoning
by: Zhang, Jiawei, et al.
Published: (2025)
by: Zhang, Jiawei, et al.
Published: (2025)
Think Just Enough: Sequence-Level Entropy as a Confidence Signal for LLM Reasoning
by: Sharma, Aman, et al.
Published: (2025)
by: Sharma, Aman, et al.
Published: (2025)
Putting the Value Back in RL: Better Test-Time Scaling by Unifying LLM Reasoners With Verifiers
by: Sareen, Kusha, et al.
Published: (2025)
by: Sareen, Kusha, et al.
Published: (2025)
Gradient Dynamics of Attention: How Cross-Entropy Sculpts Bayesian Manifolds
by: Agarwal, Naman, et al.
Published: (2025)
by: Agarwal, Naman, et al.
Published: (2025)
Entropy-Guided Data-Efficient Training for Multimodal Reasoning Reward Models
by: Yang, Shidong, et al.
Published: (2026)
by: Yang, Shidong, et al.
Published: (2026)
ProxRouter: Proximity-Weighted LLM Query Routing for Improved Robustness to Outliers
by: Patel, Shivam, et al.
Published: (2025)
by: Patel, Shivam, et al.
Published: (2025)
Fast and Effective On-policy Distillation from Reasoning Prefixes
by: Zhang, Dongxu, et al.
Published: (2026)
by: Zhang, Dongxu, et al.
Published: (2026)
Effective Exploration Based on the Structural Information Principles
by: Zeng, Xianghua, et al.
Published: (2024)
by: Zeng, Xianghua, et al.
Published: (2024)
Revisiting Entropy in Reinforcement Learning for Large Reasoning Models
by: Jin, Renren, et al.
Published: (2025)
by: Jin, Renren, et al.
Published: (2025)
YuriiFormer: A Suite of Nesterov-Accelerated Transformers
by: Zimin, Aleksandr, et al.
Published: (2026)
by: Zimin, Aleksandr, et al.
Published: (2026)
Supersonic: Learning to Generate Source Code Optimizations in C/C++
by: Chen, Zimin, et al.
Published: (2023)
by: Chen, Zimin, et al.
Published: (2023)
RetroReasoner: A Reasoning LLM for Strategic Retrosynthesis Prediction
by: Ko, Hanbum, et al.
Published: (2026)
by: Ko, Hanbum, et al.
Published: (2026)
HEAL: Hindsight Entropy-Assisted Learning for Reasoning Distillation
by: Zhang, Wenjing, et al.
Published: (2026)
by: Zhang, Wenjing, et al.
Published: (2026)
Plausibility Is Not Prediction: Contrastive Evidence for LLM-Based Cellular Perturbation Reasoning
by: Yuan, Xinyu, et al.
Published: (2026)
by: Yuan, Xinyu, et al.
Published: (2026)
RAST: Reasoning Activation in LLMs via Small-model Transfer
by: Ouyang, Siru, et al.
Published: (2025)
by: Ouyang, Siru, et al.
Published: (2025)
Generalizable Reasoning through Compositional Energy Minimization
by: Oarga, Alexandru, et al.
Published: (2025)
by: Oarga, Alexandru, et al.
Published: (2025)
On Designing Effective RL Reward at Training Time for LLM Reasoning
by: Gao, Jiaxuan, et al.
Published: (2024)
by: Gao, Jiaxuan, et al.
Published: (2024)
Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts
by: Zheng, Haizhong, et al.
Published: (2025)
by: Zheng, Haizhong, et al.
Published: (2025)
Similar Items
-
The Unreasonable Effectiveness of Eccentric Automatic Prompts
by: Battle, Rick, et al.
Published: (2024) -
Beyond Semantics: The Unreasonable Effectiveness of Reasonless Intermediate Tokens
by: Valmeekam, Karthik, et al.
Published: (2025) -
The Unreasonable Effectiveness of Model Merging for Cross-Lingual Transfer in LLMs
by: Bandarkar, Lucas, et al.
Published: (2025) -
The Unreasonable Effectiveness of Easy Training Data for Hard Tasks
by: Hase, Peter, et al.
Published: (2024) -
The Unreasonable Effectiveness of Open Science in AI: A Replication Study
by: Gundersen, Odd Erik, et al.
Published: (2024)