Learning a Continue-Thinking Token for Enhanced Test-Time Scaling
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ringel, Liran, Tolochinsky, Elad, Romano, Yaniv |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Semi-Supervised Hypothesis Testing by Betting on Predictions
von: Tenzer, Yaniv, et al.
Veröffentlicht: (2026)
von: Tenzer, Yaniv, et al.
Veröffentlicht: (2026)
Valid Best-Model Identification for LLM Evaluation via Low-Rank Factorization
von: Tolochinsky, Elad, et al.
Veröffentlicht: (2026)
von: Tolochinsky, Elad, et al.
Veröffentlicht: (2026)
Accelerating Speculative Decoding with Block Diffusion Draft Trees
von: Ringel, Liran, et al.
Veröffentlicht: (2026)
von: Ringel, Liran, et al.
Veröffentlicht: (2026)
Early Time Classification with Accumulated Accuracy Gap Control
von: Ringel, Liran, et al.
Veröffentlicht: (2024)
von: Ringel, Liran, et al.
Veröffentlicht: (2024)
Dependency-Guided Parallel Decoding in Discrete Diffusion Language Models
von: Ringel, Liran, et al.
Veröffentlicht: (2026)
von: Ringel, Liran, et al.
Veröffentlicht: (2026)
Semi-Supervised Risk Control via Prediction-Powered Inference
von: Einbinder, Bat-Sheva, et al.
Veröffentlicht: (2024)
von: Einbinder, Bat-Sheva, et al.
Veröffentlicht: (2024)
Segment-Based Attention Masking for GPTs
von: Katz, Shahar, et al.
Veröffentlicht: (2024)
von: Katz, Shahar, et al.
Veröffentlicht: (2024)
SPINE: Token-Selective Test-Time Reinforcement Learning with Entropy-Band Regularization
von: Wu, Jianghao, et al.
Veröffentlicht: (2025)
von: Wu, Jianghao, et al.
Veröffentlicht: (2025)
REMIND: Input Loss Landscapes Reveal Residual Memorization in Post-Unlearning LLMs
von: Cohen, Liran, et al.
Veröffentlicht: (2025)
von: Cohen, Liran, et al.
Veröffentlicht: (2025)
TaTToo: Tool-Grounded Thinking PRM for Test-Time Scaling in Tabular Reasoning
von: Zou, Jiaru, et al.
Veröffentlicht: (2025)
von: Zou, Jiaru, et al.
Veröffentlicht: (2025)
Aligning Tree-Search Policies with Fixed Token Budgets in Test-Time Scaling of LLMs
von: Miyamoto, Sora, et al.
Veröffentlicht: (2026)
von: Miyamoto, Sora, et al.
Veröffentlicht: (2026)
TrimR: Verifier-based Training-Free Thinking Compression for Efficient Test-Time Scaling
von: Lin, Weizhe, et al.
Veröffentlicht: (2025)
von: Lin, Weizhe, et al.
Veröffentlicht: (2025)
Latency and Token-Aware Test-Time Compute
von: Huang, Jenny Y., et al.
Veröffentlicht: (2025)
von: Huang, Jenny Y., et al.
Veröffentlicht: (2025)
Reinforcement Learning Teachers of Test Time Scaling
von: Cetin, Edoardo, et al.
Veröffentlicht: (2025)
von: Cetin, Edoardo, et al.
Veröffentlicht: (2025)
Kinetics: Rethinking Test-Time Scaling Laws
von: Sadhukhan, Ranajoy, et al.
Veröffentlicht: (2025)
von: Sadhukhan, Ranajoy, et al.
Veröffentlicht: (2025)
Test-Time Scaling with Reflective Generative Model
von: Wang, Zixiao, et al.
Veröffentlicht: (2025)
von: Wang, Zixiao, et al.
Veröffentlicht: (2025)
Test-Time Scaling Makes Overtraining Compute-Optimal
von: Roberts, Nicholas, et al.
Veröffentlicht: (2026)
von: Roberts, Nicholas, et al.
Veröffentlicht: (2026)
Chain of LoRA: Efficient Fine-tuning of Language Models via Residual Learning
von: Xia, Wenhan, et al.
Veröffentlicht: (2024)
von: Xia, Wenhan, et al.
Veröffentlicht: (2024)
Thinking in Frames: How Visual Context and Test-Time Scaling Empower Video Reasoning
von: Li, Chengzu, et al.
Veröffentlicht: (2026)
von: Li, Chengzu, et al.
Veröffentlicht: (2026)
VerifierQ: Enhancing LLM Test Time Compute with Q-Learning-based Verifiers
von: Qi, Jianing, et al.
Veröffentlicht: (2024)
von: Qi, Jianing, et al.
Veröffentlicht: (2024)
Scaling Test-Time Compute Without Verification or RL is Suboptimal
von: Setlur, Amrith, et al.
Veröffentlicht: (2025)
von: Setlur, Amrith, et al.
Veröffentlicht: (2025)
HEART: Emotionally-Driven Test-Time Scaling of Language Models
von: Pinto, Gabriela, et al.
Veröffentlicht: (2025)
von: Pinto, Gabriela, et al.
Veröffentlicht: (2025)
A Causal World Model Underlying Next Token Prediction: Exploring GPT in a Controlled Environment
von: Rohekar, Raanan Y., et al.
Veröffentlicht: (2024)
von: Rohekar, Raanan Y., et al.
Veröffentlicht: (2024)
J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge
von: Chan, Chi-Min, et al.
Veröffentlicht: (2025)
von: Chan, Chi-Min, et al.
Veröffentlicht: (2025)
TTRL: Test-Time Reinforcement Learning
von: Zuo, Yuxin, et al.
Veröffentlicht: (2025)
von: Zuo, Yuxin, et al.
Veröffentlicht: (2025)
State Tuning: State-based Test-Time Scaling on RWKV-7
von: Xiao, Liu, et al.
Veröffentlicht: (2025)
von: Xiao, Liu, et al.
Veröffentlicht: (2025)
Strategic Scaling of Test-Time Compute: A Bandit Learning Approach
von: Zuo, Bowen, et al.
Veröffentlicht: (2025)
von: Zuo, Bowen, et al.
Veröffentlicht: (2025)
Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
von: Snell, Charlie, et al.
Veröffentlicht: (2024)
von: Snell, Charlie, et al.
Veröffentlicht: (2024)
Intent-based Prompt Calibration: Enhancing prompt optimization with synthetic boundary cases
von: Levi, Elad, et al.
Veröffentlicht: (2024)
von: Levi, Elad, et al.
Veröffentlicht: (2024)
Revisiting the Test-Time Scaling of o1-like Models: Do they Truly Possess Test-Time Scaling Capabilities?
von: Zeng, Zhiyuan, et al.
Veröffentlicht: (2025)
von: Zeng, Zhiyuan, et al.
Veröffentlicht: (2025)
MetaScale: Test-Time Scaling with Evolving Meta-Thoughts
von: Liu, Qin, et al.
Veröffentlicht: (2025)
von: Liu, Qin, et al.
Veröffentlicht: (2025)
RAM-Net: Expressive Linear Attention with Selectively Addressable Memory
von: Xiao, Kaicheng, et al.
Veröffentlicht: (2026)
von: Xiao, Kaicheng, et al.
Veröffentlicht: (2026)
Think Clearly: Improving Reasoning via Redundant Token Pruning
von: Choi, Daewon, et al.
Veröffentlicht: (2025)
von: Choi, Daewon, et al.
Veröffentlicht: (2025)
Multiplex Thinking: Reasoning via Token-wise Branch-and-Merge
von: Tang, Yao, et al.
Veröffentlicht: (2026)
von: Tang, Yao, et al.
Veröffentlicht: (2026)
Over-Tokenized Transformer: Vocabulary is Generally Worth Scaling
von: Huang, Hongzhi, et al.
Veröffentlicht: (2025)
von: Huang, Hongzhi, et al.
Veröffentlicht: (2025)
It's Not That Simple. An Analysis of Simple Test-Time Scaling
von: Wu, Guojun
Veröffentlicht: (2025)
von: Wu, Guojun
Veröffentlicht: (2025)
On the Role of Temperature Sampling in Test-Time Scaling
von: Wu, Yuheng, et al.
Veröffentlicht: (2025)
von: Wu, Yuheng, et al.
Veröffentlicht: (2025)
Crosslingual Reasoning through Test-Time Scaling
von: Yong, Zheng-Xin, et al.
Veröffentlicht: (2025)
von: Yong, Zheng-Xin, et al.
Veröffentlicht: (2025)
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators
von: Zhou, Yilun, et al.
Veröffentlicht: (2025)
von: Zhou, Yilun, et al.
Veröffentlicht: (2025)
MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
von: MiniMax, et al.
Veröffentlicht: (2025)
von: MiniMax, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Semi-Supervised Hypothesis Testing by Betting on Predictions
von: Tenzer, Yaniv, et al.
Veröffentlicht: (2026) -
Valid Best-Model Identification for LLM Evaluation via Low-Rank Factorization
von: Tolochinsky, Elad, et al.
Veröffentlicht: (2026) -
Accelerating Speculative Decoding with Block Diffusion Draft Trees
von: Ringel, Liran, et al.
Veröffentlicht: (2026) -
Early Time Classification with Accumulated Accuracy Gap Control
von: Ringel, Liran, et al.
Veröffentlicht: (2024) -
Dependency-Guided Parallel Decoding in Discrete Diffusion Language Models
von: Ringel, Liran, et al.
Veröffentlicht: (2026)