Shorten After You're Right: Lazy Length Penalties for Reasoning RL
Fuente:
arXiv
Saved in:
| Main Authors: | Yuan, Danlong, Xie, Tian, Huang, Shaohan, Gong, Zhuocheng, Zhang, Huishuai, Luo, Chong, Wei, Furu, Zhao, Dongyan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Latent Preference Coding: Aligning Large Language Models via Discrete Latent Codes
by: Gong, Zhuocheng, et al.
Published: (2025)
by: Gong, Zhuocheng, et al.
Published: (2025)
SWE-MiniSandbox: Container-Free Reinforcement Learning for Building Software Engineering Agents
by: Yuan, Danlong, et al.
Published: (2026)
by: Yuan, Danlong, et al.
Published: (2026)
ReMamba: Equip Mamba with Effective Long-Sequence Modeling
by: Yuan, Danlong, et al.
Published: (2024)
by: Yuan, Danlong, et al.
Published: (2024)
Mixture-of-Modules: Reinventing Transformers as Dynamic Assemblies of Modules
by: Gong, Zhuocheng, et al.
Published: (2024)
by: Gong, Zhuocheng, et al.
Published: (2024)
Evidence-Enhanced Triplet Generation Framework for Hallucination Alleviation in Generative Question Answering
by: Du, Haowei, et al.
Published: (2024)
by: Du, Haowei, et al.
Published: (2024)
On-Policy RL with Optimal Reward Baseline
by: Hao, Yaru, et al.
Published: (2025)
by: Hao, Yaru, et al.
Published: (2025)
Bootstrap Your Own Context Length
by: Wang, Liang, et al.
Published: (2024)
by: Wang, Liang, et al.
Published: (2024)
You're worth it
by: Suzanne Jarvis
Published: (2025)
by: Suzanne Jarvis
Published: (2025)
If You're an Aries
by: Hoffman, Elizabeth
Published: (1971)
by: Hoffman, Elizabeth
Published: (1971)
ReasVQA: Advancing VideoQA with Imperfect Reasoning Process
by: Liang, Jianxin, et al.
Published: (2025)
by: Liang, Jianxin, et al.
Published: (2025)
Adapting Large Language Models to Domains via Reading Comprehension
by: Cheng, Daixuan, et al.
Published: (2023)
by: Cheng, Daixuan, et al.
Published: (2023)
Multimodal Large Language Model is a Human-Aligned Annotator for Text-to-Image Generation
by: Wu, Xun, et al.
Published: (2024)
by: Wu, Xun, et al.
Published: (2024)
Mixture of LoRA Experts
by: Wu, Xun, et al.
Published: (2024)
by: Wu, Xun, et al.
Published: (2024)
YOYO (You're On Your Own!).
by: Young, Laura
Published: (1998)
by: Young, Laura
Published: (1998)
Teaching What You're Not
Published: (2024)
Published: (2024)
You're a Parent...You're a Teacher Too. Join the Education Team.
by: Nicolau, Siobhan, et al.
Published: (1990)
by: Nicolau, Siobhan, et al.
Published: (1990)
“You're Soviet trash!—You're a liberass!”: The political life of social slurs
by: Maria Sidorkina
Published: (2025)
by: Maria Sidorkina
Published: (2025)
You're Not from Around Here, Are You?
by: Blum, Louise A.
Published: (2023)
by: Blum, Louise A.
Published: (2023)
Are You Making What You're Worth?
by: Arnold, Ruth
Published: (1999)
by: Arnold, Ruth
Published: (1999)
Synthetic Data RL: Task Definition Is All You Need
by: Guo, Yiduo, et al.
Published: (2025)
by: Guo, Yiduo, et al.
Published: (2025)
Think Only When You Need with Large Hybrid-Reasoning Models
by: Jiang, Lingjie, et al.
Published: (2025)
by: Jiang, Lingjie, et al.
Published: (2025)
What Makes Quantization for Large Language Models Hard? An Empirical Study from the Lens of Perturbation
by: Gong, Zhuocheng, et al.
Published: (2024)
by: Gong, Zhuocheng, et al.
Published: (2024)
“You're a Nobody When You're Unemployed”: Exploring the Content of Unemployed People's Stereotype
by: Charly Marie, et al.
Published: (2025)
by: Charly Marie, et al.
Published: (2025)
xRAG: Extreme Context Compression for Retrieval-augmented Generation with One Token
by: Cheng, Xin, et al.
Published: (2024)
by: Cheng, Xin, et al.
Published: (2024)
Leash: Adaptive Length Penalty and Reward Shaping for Efficient Large Reasoning Model
by: Li, Yanhao, et al.
Published: (2025)
by: Li, Yanhao, et al.
Published: (2025)
You're Dead—So What?
by: Neely, Cherly L.
Published: (2023)
by: Neely, Cherly L.
Published: (2023)
Reasoning with Exploration: An Entropy Perspective
by: Cheng, Daixuan, et al.
Published: (2025)
by: Cheng, Daixuan, et al.
Published: (2025)
Efficient Continual Pre-training by Mitigating the Stability Gap
by: Guo, Yiduo, et al.
Published: (2024)
by: Guo, Yiduo, et al.
Published: (2024)
ProactiveVideoQA: A Comprehensive Benchmark Evaluating Proactive Interactions in Video Large Language Models
by: Wang, Yueqian, et al.
Published: (2025)
by: Wang, Yueqian, et al.
Published: (2025)
Understanding Multimodal Hallucination with Parameter-Free Representation Alignment
by: Wang, Yueqian, et al.
Published: (2024)
by: Wang, Yueqian, et al.
Published: (2024)
Fast on the Easy, Deep on the Hard: Efficient Reasoning via Powered Length Penalty
by: Ling, Zehui, et al.
Published: (2025)
by: Ling, Zehui, et al.
Published: (2025)
Textual Aesthetics in Large Language Models
by: Jiang, Lingjie, et al.
Published: (2024)
by: Jiang, Lingjie, et al.
Published: (2024)
MH-MoE: Multi-Head Mixture-of-Experts
by: Huang, Shaohan, et al.
Published: (2024)
by: Huang, Shaohan, et al.
Published: (2024)
Multi-Head Mixture-of-Experts
by: Wu, Xun, et al.
Published: (2024)
by: Wu, Xun, et al.
Published: (2024)
Borrowed Templates: Is What You're Reaching For Actually Yours?
by: Sriharan, Vaz
Published: (2026)
by: Sriharan, Vaz
Published: (2026)
Ep. 1064: Why You're Falling for Your Chatbot
by: Rosehill, Daniel, et al.
Published: (2026)
by: Rosehill, Daniel, et al.
Published: (2026)
Reward Reasoning Model
by: Guo, Jiaxin, et al.
Published: (2025)
by: Guo, Jiaxin, et al.
Published: (2025)
So You're Going to Get an Intern!
by: Boardman, Edna M.
Published: (1990)
by: Boardman, Edna M.
Published: (1990)
Know When You're Wrong: Aligning Confidence with Correctness for LLM Error Detection
by: Xiaohu, Xie, et al.
Published: (2026)
by: Xiaohu, Xie, et al.
Published: (2026)
If You’re So Ethical, Why Are You So Highly Paid?
by: Pepper, Alexander
Published: (2022)
by: Pepper, Alexander
Published: (2022)
Similar Items
-
Latent Preference Coding: Aligning Large Language Models via Discrete Latent Codes
by: Gong, Zhuocheng, et al.
Published: (2025) -
SWE-MiniSandbox: Container-Free Reinforcement Learning for Building Software Engineering Agents
by: Yuan, Danlong, et al.
Published: (2026) -
ReMamba: Equip Mamba with Effective Long-Sequence Modeling
by: Yuan, Danlong, et al.
Published: (2024) -
Mixture-of-Modules: Reinventing Transformers as Dynamic Assemblies of Modules
by: Gong, Zhuocheng, et al.
Published: (2024) -
Evidence-Enhanced Triplet Generation Framework for Hallucination Alleviation in Generative Question Answering
by: Du, Haowei, et al.
Published: (2024)