Fast on the Easy, Deep on the Hard: Efficient Reasoning via Powered Length Penalty
Fuente:
arXiv
Saved in:
| Main Authors: | Ling, Zehui, Chen, Deshu, Zhang, Hongwei, Jiao, Yifeng, Guo, Xin, Cheng, Yuan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Adaptive Reasoning Executor: A Collaborative Agent System for Efficient Reasoning
by: Ling, Zehui, et al.
Published: (2025)
by: Ling, Zehui, et al.
Published: (2025)
Stop Unnecessary Reflection: Training LRMs for Efficient Reasoning with Adaptive Reflection and Length Coordinated Penalty
by: Yu, Zewei, et al.
Published: (2026)
by: Yu, Zewei, et al.
Published: (2026)
Shorten After You're Right: Lazy Length Penalties for Reasoning RL
by: Yuan, Danlong, et al.
Published: (2025)
by: Yuan, Danlong, et al.
Published: (2025)
Curriculum Reinforcement Learning from Easy to Hard Tasks Improves LLM Reasoning
by: Parashar, Shubham, et al.
Published: (2025)
by: Parashar, Shubham, et al.
Published: (2025)
Rethinking Easy-to-Hard: Limits of Curriculum Learning in Post-Training for Deductive Reasoning
by: Mordig, Maximilian, et al.
Published: (2026)
by: Mordig, Maximilian, et al.
Published: (2026)
Emergent Misalignment is Easy, Narrow Misalignment is Hard
by: Soligo, Anna, et al.
Published: (2026)
by: Soligo, Anna, et al.
Published: (2026)
Stepwise Penalization for Length-Efficient Chain-of-Thought Reasoning
by: Li, Xintong, et al.
Published: (2026)
by: Li, Xintong, et al.
Published: (2026)
AALC: Large Language Model Efficient Reasoning via Adaptive Accuracy-Length Control
by: Li, Ruosen, et al.
Published: (2025)
by: Li, Ruosen, et al.
Published: (2025)
Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards
by: Su, Jinyan, et al.
Published: (2025)
by: Su, Jinyan, et al.
Published: (2025)
Advancing Machine-Generated Text Detection from an Easy to Hard Supervision Perspective
by: Wu, Chenwang, et al.
Published: (2025)
by: Wu, Chenwang, et al.
Published: (2025)
Optimizing Length Compression in Large Reasoning Models
by: Cheng, Zhengxiang, et al.
Published: (2025)
by: Cheng, Zhengxiang, et al.
Published: (2025)
MetaLint: Easy-to-Hard Generalization for Code Linting
by: Naik, Atharva, et al.
Published: (2025)
by: Naik, Atharva, et al.
Published: (2025)
The Unreasonable Effectiveness of Easy Training Data for Hard Tasks
by: Hase, Peter, et al.
Published: (2024)
by: Hase, Peter, et al.
Published: (2024)
Can an Easy-to-Hard Curriculum Make Reasoning Emerge in Small Language Models? Evidence from a Four-Stage Curriculum on GPT-2
by: Fu, Xiang
Published: (2025)
by: Fu, Xiang
Published: (2025)
Improving Variable-Length Generation in Diffusion Language Models via Length Regularization
by: Cheng, Zicong, et al.
Published: (2026)
by: Cheng, Zicong, et al.
Published: (2026)
EasyJudge: an Easy-to-use Tool for Comprehensive Response Evaluation of LLMs
by: Li, Yijie, et al.
Published: (2024)
by: Li, Yijie, et al.
Published: (2024)
Beyond Easy Wins: A Text Hardness-Aware Benchmark for LLM-generated Text Detection
by: Ayoobi, Navid, et al.
Published: (2025)
by: Ayoobi, Navid, et al.
Published: (2025)
AI "News" Content Farms Are Easy to Make and Hard to Detect: A Case Study in Italian
by: Puccetti, Giovanni, et al.
Published: (2024)
by: Puccetti, Giovanni, et al.
Published: (2024)
Language Models Do Hard Arithmetic Tasks Easily and Hardly Do Easy Arithmetic Tasks
by: Gambardella, Andrew, et al.
Published: (2024)
by: Gambardella, Andrew, et al.
Published: (2024)
Easy-to-Hard Generalization: Scalable Alignment Beyond Human Supervision
by: Sun, Zhiqing, et al.
Published: (2024)
by: Sun, Zhiqing, et al.
Published: (2024)
The Power of Question Translation Training in Multilingual Reasoning: Broadened Scope and Deepened Insights
by: Zhu, Wenhao, et al.
Published: (2024)
by: Zhu, Wenhao, et al.
Published: (2024)
TRiMS: Real-Time Tracking of Minimal Sufficient Length for Efficient Reasoning via RL
by: Bian, Tingcheng, et al.
Published: (2026)
by: Bian, Tingcheng, et al.
Published: (2026)
On the Step Length Confounding in LLM Reasoning Data Selection
by: Wang, Bing, et al.
Published: (2026)
by: Wang, Bing, et al.
Published: (2026)
R-PRM: Reasoning-Driven Process Reward Modeling
by: She, Shuaijie, et al.
Published: (2025)
by: She, Shuaijie, et al.
Published: (2025)
ThinkLess: A Training-Free Inference-Efficient Method for Reducing Reasoning Redundancy
by: Li, Gengyang, et al.
Published: (2025)
by: Li, Gengyang, et al.
Published: (2025)
SmartThinker: Learning to Compress and Preserve Reasoning by Step-Level Length Control
by: He, Xingyang, et al.
Published: (2025)
by: He, Xingyang, et al.
Published: (2025)
SimpleDeepSearcher: Deep Information Seeking via Web-Powered Reasoning Trajectory Synthesis
by: Sun, Shuang, et al.
Published: (2025)
by: Sun, Shuang, et al.
Published: (2025)
Less is More: Compact Clue Selection for Efficient Retrieval-Augmented Generation Reasoning
by: Zhang, Qianchi, et al.
Published: (2025)
by: Zhang, Qianchi, et al.
Published: (2025)
Efficient Pretraining Length Scaling
by: Wu, Bohong, et al.
Published: (2025)
by: Wu, Bohong, et al.
Published: (2025)
A Penalty Goes a Long Way: Measuring Lexical Diversity in Synthetic Texts Under Prompt-Influenced Length Variations
by: Deshpande, Vijeta, et al.
Published: (2025)
by: Deshpande, Vijeta, et al.
Published: (2025)
The Imitation Game: Turing Machine Imitator is Length Generalizable Reasoner
by: Hua, Zhouqi, et al.
Published: (2025)
by: Hua, Zhouqi, et al.
Published: (2025)
UnsafeChain: Enhancing Reasoning Model Safety via Hard Cases
by: Tomar, Raj Vardhan, et al.
Published: (2025)
by: Tomar, Raj Vardhan, et al.
Published: (2025)
Learn to Reason Efficiently with Adaptive Length-based Reward Shaping
by: Liu, Wei, et al.
Published: (2025)
by: Liu, Wei, et al.
Published: (2025)
ING-VP: MLLMs cannot Play Easy Vision-based Games Yet
by: Zhang, Haoran, et al.
Published: (2024)
by: Zhang, Haoran, et al.
Published: (2024)
Anti-Length Shift: Dynamic Outlier Truncation for Training Efficient Reasoning Models
by: Wu, Wei, et al.
Published: (2026)
by: Wu, Wei, et al.
Published: (2026)
Planning-Aware Code Infilling via Horizon-Length Prediction
by: Ding, Yifeng, et al.
Published: (2024)
by: Ding, Yifeng, et al.
Published: (2024)
Hy-MT2: A Family of Fast, Efficient and Powerful Multilingual Translation Models in the Wild
by: Zheng, Mao, et al.
Published: (2026)
by: Zheng, Mao, et al.
Published: (2026)
SmartThinker: Progressive Chain-of-Thought Length Calibration for Efficient Large Language Model Reasoning
by: Hu, Chenzhi, et al.
Published: (2026)
by: Hu, Chenzhi, et al.
Published: (2026)
It's Not Easy Being Wrong: Large Language Models Struggle with Process of Elimination Reasoning
by: Balepur, Nishant, et al.
Published: (2023)
by: Balepur, Nishant, et al.
Published: (2023)
Preference Optimization for Reasoning with Pseudo Feedback
by: Jiao, Fangkai, et al.
Published: (2024)
by: Jiao, Fangkai, et al.
Published: (2024)
Similar Items
-
Adaptive Reasoning Executor: A Collaborative Agent System for Efficient Reasoning
by: Ling, Zehui, et al.
Published: (2025) -
Stop Unnecessary Reflection: Training LRMs for Efficient Reasoning with Adaptive Reflection and Length Coordinated Penalty
by: Yu, Zewei, et al.
Published: (2026) -
Shorten After You're Right: Lazy Length Penalties for Reasoning RL
by: Yuan, Danlong, et al.
Published: (2025) -
Curriculum Reinforcement Learning from Easy to Hard Tasks Improves LLM Reasoning
by: Parashar, Shubham, et al.
Published: (2025) -
Rethinking Easy-to-Hard: Limits of Curriculum Learning in Post-Training for Deductive Reasoning
by: Mordig, Maximilian, et al.
Published: (2026)