On the Optimal Reasoning Length for RL-Trained Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Nohara, Daisuke, Nakamura, Taishi, Yokota, Rio |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Optimal Sparsity of Mixture-of-Experts Language Models for Reasoning Tasks
von: Nakamura, Taishi, et al.
Veröffentlicht: (2025)
von: Nakamura, Taishi, et al.
Veröffentlicht: (2025)
Drop-Upcycling: Training Sparse Mixture of Experts with Partial Re-initialization
von: Nakamura, Taishi, et al.
Veröffentlicht: (2025)
von: Nakamura, Taishi, et al.
Veröffentlicht: (2025)
Reaching Beyond the Mode: RL for Distributional Reasoning in Language Models
von: Puri, Isha, et al.
Veröffentlicht: (2026)
von: Puri, Isha, et al.
Veröffentlicht: (2026)
On Designing Effective RL Reward at Training Time for LLM Reasoning
von: Gao, Jiaxuan, et al.
Veröffentlicht: (2024)
von: Gao, Jiaxuan, et al.
Veröffentlicht: (2024)
Learning to Reason as Action Abstractions with Scalable Mid-Training RL
von: Zhang, Shenao, et al.
Veröffentlicht: (2025)
von: Zhang, Shenao, et al.
Veröffentlicht: (2025)
$Q\sharp$: Provably Optimal Distributional RL for LLM Post-Training
von: Zhou, Jin Peng, et al.
Veröffentlicht: (2025)
von: Zhou, Jin Peng, et al.
Veröffentlicht: (2025)
Training Optimal Large Diffusion Language Models
von: Ni, Jinjie, et al.
Veröffentlicht: (2025)
von: Ni, Jinjie, et al.
Veröffentlicht: (2025)
Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training
von: Liu, Mingjie, et al.
Veröffentlicht: (2025)
von: Liu, Mingjie, et al.
Veröffentlicht: (2025)
ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL
von: Zhou, Yifei, et al.
Veröffentlicht: (2024)
von: Zhou, Yifei, et al.
Veröffentlicht: (2024)
RL Tango: Reinforcing Generator and Verifier Together for Language Reasoning
von: Zha, Kaiwen, et al.
Veröffentlicht: (2025)
von: Zha, Kaiwen, et al.
Veröffentlicht: (2025)
Length-MAX Tokenizer for Language Models
von: Dong, Dong, et al.
Veröffentlicht: (2025)
von: Dong, Dong, et al.
Veröffentlicht: (2025)
Clip Your Sequences Fairly: Enforcing Length Fairness for Sequence-Level RL
von: Mao, Hanyi, et al.
Veröffentlicht: (2025)
von: Mao, Hanyi, et al.
Veröffentlicht: (2025)
Exploring RL-based LLM Training for Formal Language Tasks with Programmed Rewards
von: Padula, Alexander G., et al.
Veröffentlicht: (2024)
von: Padula, Alexander G., et al.
Veröffentlicht: (2024)
Self-Training Elicits Concise Reasoning in Large Language Models
von: Munkhbat, Tergel, et al.
Veröffentlicht: (2025)
von: Munkhbat, Tergel, et al.
Veröffentlicht: (2025)
Continual Pre-Training for Cross-Lingual LLM Adaptation: Enhancing Japanese Language Capabilities
von: Fujii, Kazuki, et al.
Veröffentlicht: (2024)
von: Fujii, Kazuki, et al.
Veröffentlicht: (2024)
Training Large Language Models To Reason In Parallel With Global Forking Tokens
von: Jia, Sheng, et al.
Veröffentlicht: (2025)
von: Jia, Sheng, et al.
Veröffentlicht: (2025)
Reinforcement Learning for Reasoning in Large Language Models with One Training Example
von: Wang, Yiping, et al.
Veröffentlicht: (2025)
von: Wang, Yiping, et al.
Veröffentlicht: (2025)
GUI-Libra: Training Native GUI Agents to Reason and Act with Action-aware Supervision and Partially Verifiable RL
von: Yang, Rui, et al.
Veröffentlicht: (2026)
von: Yang, Rui, et al.
Veröffentlicht: (2026)
More Thinking, More Bias: Length-Driven Position Bias in Reasoning Models
von: Wang, Xiao
Veröffentlicht: (2026)
von: Wang, Xiao
Veröffentlicht: (2026)
GLIDE-RL: Grounded Language Instruction through DEmonstration in RL
von: Kharyal, Chaitanya, et al.
Veröffentlicht: (2024)
von: Kharyal, Chaitanya, et al.
Veröffentlicht: (2024)
VRPO: Rethinking Value Modeling for Robust RL Training under Noisy Supervision
von: Zhu, Dingwei, et al.
Veröffentlicht: (2025)
von: Zhu, Dingwei, et al.
Veröffentlicht: (2025)
Training Large Language Models for Reasoning through Reverse Curriculum Reinforcement Learning
von: Xi, Zhiheng, et al.
Veröffentlicht: (2024)
von: Xi, Zhiheng, et al.
Veröffentlicht: (2024)
FlowRL: Matching Reward Distributions for LLM Reasoning
von: Zhu, Xuekai, et al.
Veröffentlicht: (2025)
von: Zhu, Xuekai, et al.
Veröffentlicht: (2025)
Stepwise Penalization for Length-Efficient Chain-of-Thought Reasoning
von: Li, Xintong, et al.
Veröffentlicht: (2026)
von: Li, Xintong, et al.
Veröffentlicht: (2026)
Improving LoRA with Variational Learning
von: Cong, Bai, et al.
Veröffentlicht: (2025)
von: Cong, Bai, et al.
Veröffentlicht: (2025)
CoLD: Counterfactually-Guided Length Debiasing for Process Reward Models in Mathematical Reasoning
von: Zheng, Congmin, et al.
Veröffentlicht: (2025)
von: Zheng, Congmin, et al.
Veröffentlicht: (2025)
Confidence Regularized Masked Language Modeling using Text Length
von: Ji, Seunghyun, et al.
Veröffentlicht: (2025)
von: Ji, Seunghyun, et al.
Veröffentlicht: (2025)
HAPO: Training Language Models to Reason Concisely via History-Aware Policy Optimization
von: Huang, Chengyu, et al.
Veröffentlicht: (2025)
von: Huang, Chengyu, et al.
Veröffentlicht: (2025)
SFT-then-RL Outperforms Mixed-Policy Methods for LLM Reasoning
von: Limozin, Alexis, et al.
Veröffentlicht: (2026)
von: Limozin, Alexis, et al.
Veröffentlicht: (2026)
Learn to Reason Efficiently with Adaptive Length-based Reward Shaping
von: Liu, Wei, et al.
Veröffentlicht: (2025)
von: Liu, Wei, et al.
Veröffentlicht: (2025)
Distributional Clarity: The Hidden Driver of RL-Friendliness in Large Language Models
von: Sun, Shaoning, et al.
Veröffentlicht: (2026)
von: Sun, Shaoning, et al.
Veröffentlicht: (2026)
On the Expressiveness and Length Generalization of Selective State-Space Models on Regular Languages
von: Terzić, Aleksandar, et al.
Veröffentlicht: (2024)
von: Terzić, Aleksandar, et al.
Veröffentlicht: (2024)
DISPO: Enhancing Training Efficiency and Stability in Reinforcement Learning for Large Language Model Mathematical Reasoning
von: Karaman, Batuhan K., et al.
Veröffentlicht: (2026)
von: Karaman, Batuhan K., et al.
Veröffentlicht: (2026)
Open-sci-ref-0.01: open and reproducible reference baselines for language model and dataset comparison
von: Nezhurina, Marianna, et al.
Veröffentlicht: (2025)
von: Nezhurina, Marianna, et al.
Veröffentlicht: (2025)
Building a Large Japanese Web Corpus for Large Language Models
von: Okazaki, Naoaki, et al.
Veröffentlicht: (2024)
von: Okazaki, Naoaki, et al.
Veröffentlicht: (2024)
Anti-Self-Distillation for Reasoning RL via Pointwise Mutual Information
von: Shen, Guobin, et al.
Veröffentlicht: (2026)
von: Shen, Guobin, et al.
Veröffentlicht: (2026)
Synthetic Data Generation & Multi-Step RL for Reasoning & Tool Use
von: Goldie, Anna, et al.
Veröffentlicht: (2025)
von: Goldie, Anna, et al.
Veröffentlicht: (2025)
Variational Low-Rank Adaptation Using IVON
von: Cong, Bai, et al.
Veröffentlicht: (2024)
von: Cong, Bai, et al.
Veröffentlicht: (2024)
Asynchronous RLHF: Faster and More Efficient Off-Policy RL for Language Models
von: Noukhovitch, Michael, et al.
Veröffentlicht: (2024)
von: Noukhovitch, Michael, et al.
Veröffentlicht: (2024)
Evaluating & Reducing Deceptive Dialogue From Language Models with Multi-turn RL
von: Abdulhai, Marwa, et al.
Veröffentlicht: (2025)
von: Abdulhai, Marwa, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Optimal Sparsity of Mixture-of-Experts Language Models for Reasoning Tasks
von: Nakamura, Taishi, et al.
Veröffentlicht: (2025) -
Drop-Upcycling: Training Sparse Mixture of Experts with Partial Re-initialization
von: Nakamura, Taishi, et al.
Veröffentlicht: (2025) -
Reaching Beyond the Mode: RL for Distributional Reasoning in Language Models
von: Puri, Isha, et al.
Veröffentlicht: (2026) -
On Designing Effective RL Reward at Training Time for LLM Reasoning
von: Gao, Jiaxuan, et al.
Veröffentlicht: (2024) -
Learning to Reason as Action Abstractions with Scalable Mid-Training RL
von: Zhang, Shenao, et al.
Veröffentlicht: (2025)