Rethinking Fine-Tuning when Scaling Test-Time Compute: Limiting Confidence Improves Mathematical Reasoning
Fuente:
arXiv
Salvato in:
| Autori principali: | Chen, Feng, Raventos, Allan, Cheng, Nan, Ganguli, Surya, Druckmann, Shaul |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Deriving Neural Scaling Laws from the statistics of natural language
di: Cagnetta, Francesco, et al.
Pubblicazione: (2026)
di: Cagnetta, Francesco, et al.
Pubblicazione: (2026)
Unlocking LLM Creativity in Science through Analogical Reasoning
di: Shen, Andrew, et al.
Pubblicazione: (2026)
di: Shen, Andrew, et al.
Pubblicazione: (2026)
Get rich quick: exact solutions reveal how unbalanced initializations promote rapid feature learning
di: Kunin, Daniel, et al.
Pubblicazione: (2024)
di: Kunin, Daniel, et al.
Pubblicazione: (2024)
Informed Correctors for Discrete Diffusion Models
di: Zhao, Yixiu, et al.
Pubblicazione: (2024)
di: Zhao, Yixiu, et al.
Pubblicazione: (2024)
Stochastic Collapse: How Gradient Noise Attracts SGD Dynamics Towards Simpler Subnetworks
di: Chen, Feng, et al.
Pubblicazione: (2023)
di: Chen, Feng, et al.
Pubblicazione: (2023)
Fooling LLM graders into giving better grades through neural activity guided adversarial prompting
di: Yamamura, Atsushi, et al.
Pubblicazione: (2024)
di: Yamamura, Atsushi, et al.
Pubblicazione: (2024)
Multi-Agent Reasoning Improves Compute Efficiency: Pareto-Optimal Test-Time Scaling
di: Wunderlich, Florian Valentin, et al.
Pubblicazione: (2026)
di: Wunderlich, Florian Valentin, et al.
Pubblicazione: (2026)
Rethinking Optimal Verification Granularity for Compute-Efficient Test-Time Scaling
di: Chen, Hao Mark, et al.
Pubblicazione: (2025)
di: Chen, Hao Mark, et al.
Pubblicazione: (2025)
A Cross-Modal Approach to Silent Speech with LLM-Enhanced Recognition
di: Benster, Tyler, et al.
Pubblicazione: (2024)
di: Benster, Tyler, et al.
Pubblicazione: (2024)
Examining False Positives under Inference Scaling for Mathematical Reasoning
di: Wang, Yu, et al.
Pubblicazione: (2025)
di: Wang, Yu, et al.
Pubblicazione: (2025)
BitCal-TTS: Bit-Calibrated Test-Time Scaling for Quantized Reasoning Models
di: Patarlapalli, Sai Babu, et al.
Pubblicazione: (2026)
di: Patarlapalli, Sai Babu, et al.
Pubblicazione: (2026)
MathSE: Improving Multimodal Mathematical Reasoning via Self-Evolving Iterative Reflection and Reward-Guided Fine-Tuning
di: Chen, Jinhao, et al.
Pubblicazione: (2025)
di: Chen, Jinhao, et al.
Pubblicazione: (2025)
MathScale: Scaling Instruction Tuning for Mathematical Reasoning
di: Tang, Zhengyang, et al.
Pubblicazione: (2024)
di: Tang, Zhengyang, et al.
Pubblicazione: (2024)
TMAS: Scaling Test-Time Compute via Multi-Agent Synergy
di: Wu, George, et al.
Pubblicazione: (2026)
di: Wu, George, et al.
Pubblicazione: (2026)
Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning
di: Qu, Yuxiao, et al.
Pubblicazione: (2025)
di: Qu, Yuxiao, et al.
Pubblicazione: (2025)
An analytic theory of creativity in convolutional diffusion models
di: Kamb, Mason, et al.
Pubblicazione: (2024)
di: Kamb, Mason, et al.
Pubblicazione: (2024)
Rethinking Reward Models for Multi-Domain Test-Time Scaling
di: Lee, Dong Bok, et al.
Pubblicazione: (2025)
di: Lee, Dong Bok, et al.
Pubblicazione: (2025)
Maximizing Confidence Alone Improves Reasoning
di: Prabhudesai, Mihir, et al.
Pubblicazione: (2025)
di: Prabhudesai, Mihir, et al.
Pubblicazione: (2025)
NeuroProlog: Multi-Task Fine-Tuning for Neurosymbolic Mathematical Reasoning via the Cocktail Effect
di: Zunjare, Pratibha, et al.
Pubblicazione: (2026)
di: Zunjare, Pratibha, et al.
Pubblicazione: (2026)
CoRefine: Confidence-Guided Self-Refinement for Adaptive Test-Time Compute
di: Jin, Chen, et al.
Pubblicazione: (2026)
di: Jin, Chen, et al.
Pubblicazione: (2026)
ReThinker: Scientific Reasoning by Rethinking with Guided Reflection and Confidence Control
di: Tang, Zhentao, et al.
Pubblicazione: (2026)
di: Tang, Zhentao, et al.
Pubblicazione: (2026)
Log-Augmented Generation: Scaling Test-Time Reasoning with Reusable Computation
di: Chen, Peter Baile, et al.
Pubblicazione: (2025)
di: Chen, Peter Baile, et al.
Pubblicazione: (2025)
TFGN: Task-Free, Replay-Free Continual Pre-Training Without Catastrophic Forgetting at LLM Scale
di: Ganguli, Anurup
Pubblicazione: (2026)
di: Ganguli, Anurup
Pubblicazione: (2026)
Towards Thinking-Optimal Scaling of Test-Time Compute for LLM Reasoning
di: Yang, Wenkai, et al.
Pubblicazione: (2025)
di: Yang, Wenkai, et al.
Pubblicazione: (2025)
Forest-of-Thought: Scaling Test-Time Compute for Enhancing LLM Reasoning
di: Bi, Zhenni, et al.
Pubblicazione: (2024)
di: Bi, Zhenni, et al.
Pubblicazione: (2024)
Less is More: Improving LLM Reasoning with Minimal Test-Time Intervention
di: Yang, Zhen, et al.
Pubblicazione: (2025)
di: Yang, Zhen, et al.
Pubblicazione: (2025)
Inverse Scaling in Test-Time Compute
di: Gema, Aryo Pradipta, et al.
Pubblicazione: (2025)
di: Gema, Aryo Pradipta, et al.
Pubblicazione: (2025)
Rethinking Supervised Fine-Tuning: Emphasizing Key Answer Tokens for Improved LLM Accuracy
di: Shi, Xiaofeng, et al.
Pubblicazione: (2025)
di: Shi, Xiaofeng, et al.
Pubblicazione: (2025)
Guided by Gut: Efficient Test-Time Scaling with Reinforced Intrinsic Confidence
di: Ghasemabadi, Amirhosein, et al.
Pubblicazione: (2025)
di: Ghasemabadi, Amirhosein, et al.
Pubblicazione: (2025)
BrowseConf: Confidence-Guided Test-Time Scaling for Web Agents
di: Ou, Litu, et al.
Pubblicazione: (2025)
di: Ou, Litu, et al.
Pubblicazione: (2025)
On-Policy Supervised Fine-Tuning for Efficient Reasoning
di: Zhao, Anhao, et al.
Pubblicazione: (2026)
di: Zhao, Anhao, et al.
Pubblicazione: (2026)
Mitigating Strategy-Selection Bias in Reasoning for More Effective Test-Time Scaling
di: Wu, Zongqian, et al.
Pubblicazione: (2025)
di: Wu, Zongqian, et al.
Pubblicazione: (2025)
Rethinking the Role of Prompting Strategies in LLM Test-Time Scaling: A Perspective of Probability Theory
di: Liu, Yexiang, et al.
Pubblicazione: (2025)
di: Liu, Yexiang, et al.
Pubblicazione: (2025)
Beyond Model Scaling: Test-Time Intervention for Efficient Deep Reasoning
di: Wang, Qianyue, et al.
Pubblicazione: (2026)
di: Wang, Qianyue, et al.
Pubblicazione: (2026)
Efficiently Learning at Test-Time: Active Fine-Tuning of LLMs
di: Hübotter, Jonas, et al.
Pubblicazione: (2024)
di: Hübotter, Jonas, et al.
Pubblicazione: (2024)
Agentic Test-Time Scaling for WebAgents
di: Lee, Nicholas, et al.
Pubblicazione: (2026)
di: Lee, Nicholas, et al.
Pubblicazione: (2026)
MatryoshkaThinking: Recursive Test-Time Scaling Enables Efficient Reasoning
di: Chen, Hongwei, et al.
Pubblicazione: (2025)
di: Chen, Hongwei, et al.
Pubblicazione: (2025)
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction
di: Shen, Junhong, et al.
Pubblicazione: (2025)
di: Shen, Junhong, et al.
Pubblicazione: (2025)
GRIP: In-Parameter Graph Reasoning through Fine-Tuning Large Language Models
di: Feng, Jiarui, et al.
Pubblicazione: (2025)
di: Feng, Jiarui, et al.
Pubblicazione: (2025)
ScaleRTL: Scaling LLMs with Reasoning Data and Test-Time Compute for Accurate RTL Code Generation
di: Deng, Chenhui, et al.
Pubblicazione: (2025)
di: Deng, Chenhui, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Deriving Neural Scaling Laws from the statistics of natural language
di: Cagnetta, Francesco, et al.
Pubblicazione: (2026) -
Unlocking LLM Creativity in Science through Analogical Reasoning
di: Shen, Andrew, et al.
Pubblicazione: (2026) -
Get rich quick: exact solutions reveal how unbalanced initializations promote rapid feature learning
di: Kunin, Daniel, et al.
Pubblicazione: (2024) -
Informed Correctors for Discrete Diffusion Models
di: Zhao, Yixiu, et al.
Pubblicazione: (2024) -
Stochastic Collapse: How Gradient Noise Attracts SGD Dynamics Towards Simpler Subnetworks
di: Chen, Feng, et al.
Pubblicazione: (2023)