Thinking Long, but Short: Stable Sequential Test-Time Scaling for Large Reasoning Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Metel, Michael R., Cui, Yufei, Chen, Boxing, Parthasarathi, Prasanna |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Nested-ReFT: Efficient Reinforcement Learning for Large Language Model Fine-Tuning via Off-Policy Rollouts
von: Heuillet, Maxime, et al.
Veröffentlicht: (2025)
von: Heuillet, Maxime, et al.
Veröffentlicht: (2025)
Do Robot Snakes Dream like Electric Sheep? Investigating the Effects of Architectural Inductive Biases on Hallucination
von: Huang, Jerry, et al.
Veröffentlicht: (2024)
von: Huang, Jerry, et al.
Veröffentlicht: (2024)
GRPO-$λ$: Credit Assignment improves LLM Reasoning
von: Parthasarathi, Prasanna, et al.
Veröffentlicht: (2025)
von: Parthasarathi, Prasanna, et al.
Veröffentlicht: (2025)
MatryoshkaThinking: Recursive Test-Time Scaling Enables Efficient Reasoning
von: Chen, Hongwei, et al.
Veröffentlicht: (2025)
von: Chen, Hongwei, et al.
Veröffentlicht: (2025)
Does Thinking More always Help? Mirage of Test-Time Scaling in Reasoning Models
von: Ghosal, Soumya Suvra, et al.
Veröffentlicht: (2025)
von: Ghosal, Soumya Suvra, et al.
Veröffentlicht: (2025)
Towards Thinking-Optimal Scaling of Test-Time Compute for LLM Reasoning
von: Yang, Wenkai, et al.
Veröffentlicht: (2025)
von: Yang, Wenkai, et al.
Veröffentlicht: (2025)
Do Large Language Models Know How Much They Know?
von: Prato, Gabriele, et al.
Veröffentlicht: (2025)
von: Prato, Gabriele, et al.
Veröffentlicht: (2025)
Batch-Max: Higher LLM Throughput using Larger Batch Sizes and KV Cache Compression
von: Metel, Michael R., et al.
Veröffentlicht: (2024)
von: Metel, Michael R., et al.
Veröffentlicht: (2024)
Beyond Mode-Seeking RL: Trajectory-Balance Post-Training for Diffusion Language Models
von: Ahmadi, Saba, et al.
Veröffentlicht: (2026)
von: Ahmadi, Saba, et al.
Veröffentlicht: (2026)
Train Long, Think Short: Curriculum Learning for Efficient Reasoning
von: Hammoud, Hasan Abed Al Kader, et al.
Veröffentlicht: (2025)
von: Hammoud, Hasan Abed Al Kader, et al.
Veröffentlicht: (2025)
LongReasonArena: A Long Reasoning Benchmark for Large Language Models
von: Ding, Jiayu, et al.
Veröffentlicht: (2025)
von: Ding, Jiayu, et al.
Veröffentlicht: (2025)
Towards Practical Tool Usage for Continually Learning LLMs
von: Huang, Jerry, et al.
Veröffentlicht: (2024)
von: Huang, Jerry, et al.
Veröffentlicht: (2024)
Scaling over Scaling: Exploring Test-Time Scaling Plateau in Large Reasoning Models
von: Wang, Jian, et al.
Veröffentlicht: (2025)
von: Wang, Jian, et al.
Veröffentlicht: (2025)
InftyThink: Breaking the Length Limits of Long-Context Reasoning in Large Language Models
von: Yan, Yuchen, et al.
Veröffentlicht: (2025)
von: Yan, Yuchen, et al.
Veröffentlicht: (2025)
m1: Unleash the Potential of Test-Time Scaling for Medical Reasoning with Large Language Models
von: Huang, Xiaoke, et al.
Veröffentlicht: (2025)
von: Huang, Xiaoke, et al.
Veröffentlicht: (2025)
Modeling Hierarchical Thinking in Large Reasoning Models
von: Shahariar, G M, et al.
Veröffentlicht: (2025)
von: Shahariar, G M, et al.
Veröffentlicht: (2025)
A Survey on Test-Time Scaling in Large Language Models: What, How, Where, and How Well?
von: Zhang, Qiyuan, et al.
Veröffentlicht: (2025)
von: Zhang, Qiyuan, et al.
Veröffentlicht: (2025)
TaTToo: Tool-Grounded Thinking PRM for Test-Time Scaling in Tabular Reasoning
von: Zou, Jiaru, et al.
Veröffentlicht: (2025)
von: Zou, Jiaru, et al.
Veröffentlicht: (2025)
Logical Reasoning with Outcome Reward Models for Test-Time Scaling
von: Thatikonda, Ramya Keerthy, et al.
Veröffentlicht: (2025)
von: Thatikonda, Ramya Keerthy, et al.
Veröffentlicht: (2025)
Speculative Thinking: Enhancing Small-Model Reasoning with Large Model Guidance at Inference Time
von: Yang, Wang, et al.
Veröffentlicht: (2025)
von: Yang, Wang, et al.
Veröffentlicht: (2025)
To Think or Not To Think, That is The Question for Large Reasoning Models in Theory of Mind Tasks
von: Gong, Nanxu, et al.
Veröffentlicht: (2026)
von: Gong, Nanxu, et al.
Veröffentlicht: (2026)
Precedent-Informed Reasoning: Mitigating Overthinking in Large Reasoning Models via Test-Time Precedent Learning
von: Wang, Qianyue, et al.
Veröffentlicht: (2026)
von: Wang, Qianyue, et al.
Veröffentlicht: (2026)
Think Again! The Effect of Test-Time Compute on Preferences, Opinions, and Beliefs of Large Language Models
von: Kour, George, et al.
Veröffentlicht: (2025)
von: Kour, George, et al.
Veröffentlicht: (2025)
Thinking on the Fly: Test-Time Reasoning Enhancement via Latent Thought Policy Optimization
von: Ye, Wengao, et al.
Veröffentlicht: (2025)
von: Ye, Wengao, et al.
Veröffentlicht: (2025)
Lissard: Long and Simple Sequential Reasoning Datasets
von: Bueno, Mirelle, et al.
Veröffentlicht: (2024)
von: Bueno, Mirelle, et al.
Veröffentlicht: (2024)
Dynamic Thinking-Token Selection for Efficient Reasoning in Large Reasoning Models
von: Guo, Zhenyuan, et al.
Veröffentlicht: (2026)
von: Guo, Zhenyuan, et al.
Veröffentlicht: (2026)
Parallel Test-Time Scaling for Latent Reasoning Models
von: You, Runyang, et al.
Veröffentlicht: (2025)
von: You, Runyang, et al.
Veröffentlicht: (2025)
ReProbe: Efficient Test-Time Scaling of Multi-Step Reasoning by Probing Internal States of Large Language Models
von: Ni, Jingwei, et al.
Veröffentlicht: (2025)
von: Ni, Jingwei, et al.
Veröffentlicht: (2025)
Think Before Recommend: Unleashing the Latent Reasoning Power for Sequential Recommendation
von: Tang, Jiakai, et al.
Veröffentlicht: (2025)
von: Tang, Jiakai, et al.
Veröffentlicht: (2025)
Controlling Thinking Speed in Reasoning Models
von: Lin, Zhengkai, et al.
Veröffentlicht: (2025)
von: Lin, Zhengkai, et al.
Veröffentlicht: (2025)
Resona: Improving Context Copying in Linear Recurrence Models with Retrieval
von: Wang, Xinyu, et al.
Veröffentlicht: (2025)
von: Wang, Xinyu, et al.
Veröffentlicht: (2025)
Think$^{2}$: Grounded Metacognitive Reasoning in Large Language Models
von: Elenjical, Abraham Paul, et al.
Veröffentlicht: (2026)
von: Elenjical, Abraham Paul, et al.
Veröffentlicht: (2026)
Exploring the System 1 Thinking Capability of Large Reasoning Models
von: Zhang, Wenyuan, et al.
Veröffentlicht: (2025)
von: Zhang, Wenyuan, et al.
Veröffentlicht: (2025)
Bridging the Reasoning Gap in Vietnamese with Small Language Models via Test-Time Scaling
von: Trung, Bui The, et al.
Veröffentlicht: (2026)
von: Trung, Bui The, et al.
Veröffentlicht: (2026)
To Think or Not to Think: Exploring the Unthinking Vulnerability in Large Reasoning Models
von: Zhu, Zihao, et al.
Veröffentlicht: (2025)
von: Zhu, Zihao, et al.
Veröffentlicht: (2025)
Log-Augmented Generation: Scaling Test-Time Reasoning with Reusable Computation
von: Chen, Peter Baile, et al.
Veröffentlicht: (2025)
von: Chen, Peter Baile, et al.
Veröffentlicht: (2025)
Thinking in Frames: How Visual Context and Test-Time Scaling Empower Video Reasoning
von: Li, Chengzu, et al.
Veröffentlicht: (2026)
von: Li, Chengzu, et al.
Veröffentlicht: (2026)
LSTPrompt: Large Language Models as Zero-Shot Time Series Forecasters by Long-Short-Term Prompting
von: Liu, Haoxin, et al.
Veröffentlicht: (2024)
von: Liu, Haoxin, et al.
Veröffentlicht: (2024)
Do Thinking Tokens Help or Trap? Towards More Efficient Large Reasoning Model
von: Ding, Bowen, et al.
Veröffentlicht: (2025)
von: Ding, Bowen, et al.
Veröffentlicht: (2025)
ThinkPilot: Steering Reasoning Models via Automated Think-prefixes Optimization
von: Li, Sunzhu, et al.
Veröffentlicht: (2025)
von: Li, Sunzhu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Nested-ReFT: Efficient Reinforcement Learning for Large Language Model Fine-Tuning via Off-Policy Rollouts
von: Heuillet, Maxime, et al.
Veröffentlicht: (2025) -
Do Robot Snakes Dream like Electric Sheep? Investigating the Effects of Architectural Inductive Biases on Hallucination
von: Huang, Jerry, et al.
Veröffentlicht: (2024) -
GRPO-$λ$: Credit Assignment improves LLM Reasoning
von: Parthasarathi, Prasanna, et al.
Veröffentlicht: (2025) -
MatryoshkaThinking: Recursive Test-Time Scaling Enables Efficient Reasoning
von: Chen, Hongwei, et al.
Veröffentlicht: (2025) -
Does Thinking More always Help? Mirage of Test-Time Scaling in Reasoning Models
von: Ghosal, Soumya Suvra, et al.
Veröffentlicht: (2025)