Maximizing Prefix-Confidence at Test-Time Efficiently Improves Mathematical Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Otth, Matthias, Hübotter, Jonas, Hakimi, Ido, Krause, Andreas |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Efficiently Learning at Test-Time: Active Fine-Tuning of LLMs
by: Hübotter, Jonas, et al.
Published: (2024)
by: Hübotter, Jonas, et al.
Published: (2024)
Local Mixtures of Experts: Essentially Free Test-Time Training via Model Merging
by: Bertolissi, Ryo, et al.
Published: (2025)
by: Bertolissi, Ryo, et al.
Published: (2025)
Learning on the Job: Test-Time Curricula for Targeted Reinforcement Learning
by: Hübotter, Jonas, et al.
Published: (2025)
by: Hübotter, Jonas, et al.
Published: (2025)
Majority Voting for Code Generation
by: Launer, Tim, et al.
Published: (2026)
by: Launer, Tim, et al.
Published: (2026)
LITE: Efficiently Estimating Gaussian Probability of Maximality
by: Menet, Nicolas, et al.
Published: (2025)
by: Menet, Nicolas, et al.
Published: (2025)
Probabilistic Artificial Intelligence
by: Krause, Andreas, et al.
Published: (2025)
by: Krause, Andreas, et al.
Published: (2025)
Test-time Offline Reinforcement Learning on Goal-related Experience
by: Bagatella, Marco, et al.
Published: (2025)
by: Bagatella, Marco, et al.
Published: (2025)
Specialization after Generalization: Towards Understanding Test-Time Training in Foundation Models
by: Hübotter, Jonas, et al.
Published: (2025)
by: Hübotter, Jonas, et al.
Published: (2025)
Active Fine-Tuning of Multi-Task Policies
by: Bagatella, Marco, et al.
Published: (2024)
by: Bagatella, Marco, et al.
Published: (2024)
Active Few-Shot Fine-Tuning
by: Hübotter, Jonas, et al.
Published: (2024)
by: Hübotter, Jonas, et al.
Published: (2024)
Transductive Active Learning: Theory and Applications
by: Hübotter, Jonas, et al.
Published: (2024)
by: Hübotter, Jonas, et al.
Published: (2024)
DISCOVER: Automated Curricula for Sparse-Reward Reinforcement Learning
by: Diaz-Bone, Leander, et al.
Published: (2025)
by: Diaz-Bone, Leander, et al.
Published: (2025)
Reinforcement Learning via Self-Distillation
by: Hübotter, Jonas, et al.
Published: (2026)
by: Hübotter, Jonas, et al.
Published: (2026)
Maximizing Confidence Alone Improves Reasoning
by: Prabhudesai, Mihir, et al.
Published: (2025)
by: Prabhudesai, Mihir, et al.
Published: (2025)
Rethinking Fine-Tuning when Scaling Test-Time Compute: Limiting Confidence Improves Mathematical Reasoning
by: Chen, Feng, et al.
Published: (2025)
by: Chen, Feng, et al.
Published: (2025)
ActiveUltraFeedback: Efficient Preference Data Generation using Active Learning
by: Melikidze, Davit, et al.
Published: (2026)
by: Melikidze, Davit, et al.
Published: (2026)
Bias-Restrained Prefix Representation Finetuning for Mathematical Reasoning
by: Liang, Sirui, et al.
Published: (2025)
by: Liang, Sirui, et al.
Published: (2025)
Aligning Language Models from User Interactions
by: Buening, Thomas Kleine, et al.
Published: (2026)
by: Buening, Thomas Kleine, et al.
Published: (2026)
Self-Distillation Enables Continual Learning
by: Shenfeld, Idan, et al.
Published: (2026)
by: Shenfeld, Idan, et al.
Published: (2026)
Beyond Pairwise Correlations: Higher-Order Redundancies in Self-Supervised Representation Learning
by: Zollikofer, David, et al.
Published: (2024)
by: Zollikofer, David, et al.
Published: (2024)
Confidence Estimation via Sequential Likelihood Mixing
by: Kirschner, Johannes, et al.
Published: (2025)
by: Kirschner, Johannes, et al.
Published: (2025)
RewardUQ: A Unified Framework for Uncertainty-Aware Reward Models
by: Yang, Daniel, et al.
Published: (2026)
by: Yang, Daniel, et al.
Published: (2026)
Prefix Grouper: Efficient GRPO Training through Shared-Prefix Forward
by: Liu, Zikang, et al.
Published: (2025)
by: Liu, Zikang, et al.
Published: (2025)
Fast and Effective On-policy Distillation from Reasoning Prefixes
by: Zhang, Dongxu, et al.
Published: (2026)
by: Zhang, Dongxu, et al.
Published: (2026)
Sequential-Parallel Duality in Prefix Scannable Models
by: Yau, Morris, et al.
Published: (2025)
by: Yau, Morris, et al.
Published: (2025)
Confidence over Time: Confidence Calibration with Temporal Logic for Large Language Model Reasoning
by: Mao, Zhenjiang, et al.
Published: (2026)
by: Mao, Zhenjiang, et al.
Published: (2026)
Tool-R0: Self-Evolving LLM Agents for Tool-Learning from Zero Data
by: Acikgoz, Emre Can, et al.
Published: (2026)
by: Acikgoz, Emre Can, et al.
Published: (2026)
Simulation Priors for Data-Efficient Deep Learning
by: Treven, Lenart, et al.
Published: (2025)
by: Treven, Lenart, et al.
Published: (2025)
Test-Time Tuned Language Models Enable End-to-end De Novo Molecular Structure Generation from MS/MS Spectra
by: Mismetti, Laura, et al.
Published: (2025)
by: Mismetti, Laura, et al.
Published: (2025)
FOReCAst: The Future Outcome Reasoning and Confidence Assessment Benchmark
by: Yuan, Zhangdie, et al.
Published: (2025)
by: Yuan, Zhangdie, et al.
Published: (2025)
Data-Efficient Task Generalization via Probabilistic Model-based Meta Reinforcement Learning
by: Bhardwaj, Arjun, et al.
Published: (2023)
by: Bhardwaj, Arjun, et al.
Published: (2023)
Efficient Test-Time Inference via Deterministic Exploration of Truncated Decoding Trees
by: Li, Xueyan, et al.
Published: (2026)
by: Li, Xueyan, et al.
Published: (2026)
When to Sense and Control? A Time-adaptive Approach for Continuous-Time RL
by: Treven, Lenart, et al.
Published: (2024)
by: Treven, Lenart, et al.
Published: (2024)
MathTutorBench: A Benchmark for Measuring Open-ended Pedagogical Capabilities of LLM Tutors
by: Macina, Jakub, et al.
Published: (2025)
by: Macina, Jakub, et al.
Published: (2025)
CER: Confidence Enhanced Reasoning in LLMs
by: Razghandi, Ali, et al.
Published: (2025)
by: Razghandi, Ali, et al.
Published: (2025)
Form Follows Function: Recursive Stem Model
by: Hakimi, Navid
Published: (2026)
by: Hakimi, Navid
Published: (2026)
Plan and Budget: Effective and Efficient Test-Time Scaling on Reasoning Large Language Models
by: Lin, Junhong, et al.
Published: (2025)
by: Lin, Junhong, et al.
Published: (2025)
ECHO: Entropy-Confidence Hybrid Optimization for Test-Time Reinforcement Learning
by: Zhao, Chu, et al.
Published: (2026)
by: Zhao, Chu, et al.
Published: (2026)
PrefixQuant: Eliminating Outliers by Prefixed Tokens for Large Language Models Quantization
by: Chen, Mengzhao, et al.
Published: (2024)
by: Chen, Mengzhao, et al.
Published: (2024)
PERK: Long-Context Reasoning as Parameter-Efficient Test-Time Learning
by: Chen, Zeming, et al.
Published: (2025)
by: Chen, Zeming, et al.
Published: (2025)
Similar Items
-
Efficiently Learning at Test-Time: Active Fine-Tuning of LLMs
by: Hübotter, Jonas, et al.
Published: (2024) -
Local Mixtures of Experts: Essentially Free Test-Time Training via Model Merging
by: Bertolissi, Ryo, et al.
Published: (2025) -
Learning on the Job: Test-Time Curricula for Targeted Reinforcement Learning
by: Hübotter, Jonas, et al.
Published: (2025) -
Majority Voting for Code Generation
by: Launer, Tim, et al.
Published: (2026) -
LITE: Efficiently Estimating Gaussian Probability of Maximality
by: Menet, Nicolas, et al.
Published: (2025)