How Uncertainty Estimation Scales with Sampling in Reasoning Models
Fuente:
arXiv
Saved in:
| Main Authors: | Del, Maksym, Kängsepp, Markus, Domnich, Marharyta, Tampuu, Ardi, Yankovskaya, Lisa, Kull, Meelis, Fishel, Mark |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
On the Usefulness of the Fit-on-the-Test View on Evaluating Calibration of Classifiers
by: Kängsepp, Markus, et al.
Published: (2022)
by: Kängsepp, Markus, et al.
Published: (2022)
Enhancing Counterfactual Explanation Search with Diffusion Distance and Directional Coherence
by: Domnich, Marharyta, et al.
Published: (2024)
by: Domnich, Marharyta, et al.
Published: (2024)
Cross-model Fairness: Empirical Study of Fairness and Ethics Under Model Multiplicity
by: Sokol, Kacper, et al.
Published: (2022)
by: Sokol, Kacper, et al.
Published: (2022)
Cautious Calibration in Binary Classification
by: Allikivi, Mari-Liis, et al.
Published: (2024)
by: Allikivi, Mari-Liis, et al.
Published: (2024)
Towards Unifying Evaluation of Counterfactual Explanations: Leveraging Large Language Models for Human-Centric Assessments
by: Domnich, Marharyta, et al.
Published: (2024)
by: Domnich, Marharyta, et al.
Published: (2024)
Combating the effects of speed and delays in end-to-end self-driving
by: Tampuu, Ardi, et al.
Published: (2023)
by: Tampuu, Ardi, et al.
Published: (2023)
COIN: Counterfactual inpainting for weakly supervised semantic segmentation for medical images
by: Shvetsov, Dmytro, et al.
Published: (2024)
by: Shvetsov, Dmytro, et al.
Published: (2024)
TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning
by: Zhang, Tunyu, et al.
Published: (2025)
by: Zhang, Tunyu, et al.
Published: (2025)
Improving Calibration by Relating Focal Loss, Temperature Scaling, and Properness
by: Komisarenko, Viacheslav, et al.
Published: (2024)
by: Komisarenko, Viacheslav, et al.
Published: (2024)
Tracing Uncertainty in Language Model "Reasoning"
by: Grünefeld, Nils, et al.
Published: (2026)
by: Grünefeld, Nils, et al.
Published: (2026)
Incoherence as Oracle-less Measure of Error in LLM-Based Code Generation
by: Valentin, Thomas, et al.
Published: (2025)
by: Valentin, Thomas, et al.
Published: (2025)
Enhancing web traffic attacks identification through ensemble methods and feature selection
by: Urda, Daniel, et al.
Published: (2024)
by: Urda, Daniel, et al.
Published: (2024)
Revisiting Uncertainty Estimation and Calibration of Large Language Models
by: Tao, Linwei, et al.
Published: (2025)
by: Tao, Linwei, et al.
Published: (2025)
The Confidence Trap: Gender Bias and Predictive Certainty in LLMs
by: Sabir, Ahmed, et al.
Published: (2026)
by: Sabir, Ahmed, et al.
Published: (2026)
Scaling over Scaling: Exploring Test-Time Scaling Plateau in Large Reasoning Models
by: Wang, Jian, et al.
Published: (2025)
by: Wang, Jian, et al.
Published: (2025)
Perceptions of Linguistic Uncertainty by Language Models and Humans
by: Belem, Catarina G, et al.
Published: (2024)
by: Belem, Catarina G, et al.
Published: (2024)
HS-STaR: Hierarchical Sampling for Self-Taught Reasoners via Difficulty Estimation and Budget Reallocation
by: Xiong, Feng, et al.
Published: (2025)
by: Xiong, Feng, et al.
Published: (2025)
Parallel Test-Time Scaling for Latent Reasoning Models
by: You, Runyang, et al.
Published: (2025)
by: You, Runyang, et al.
Published: (2025)
Deep Language Geometry: Constructing a Metric Space from LLM Weights
by: Shamrai, Maksym, et al.
Published: (2025)
by: Shamrai, Maksym, et al.
Published: (2025)
Does Refusal Training in LLMs Generalize to the Past Tense?
by: Andriushchenko, Maksym, et al.
Published: (2024)
by: Andriushchenko, Maksym, et al.
Published: (2024)
GENUINE: Graph Enhanced Multi-level Uncertainty Estimation for Large Language Models
by: Wang, Tuo, et al.
Published: (2025)
by: Wang, Tuo, et al.
Published: (2025)
Improving Instruction Following in Language Models through Proxy-Based Uncertainty Estimation
by: Lee, JoonHo, et al.
Published: (2024)
by: Lee, JoonHo, et al.
Published: (2024)
Reasoning with Sampling: Your Base Model is Smarter Than You Think
by: Karan, Aayush, et al.
Published: (2025)
by: Karan, Aayush, et al.
Published: (2025)
Estimation of Concept Explanations Should be Uncertainty Aware
by: Piratla, Vihari, et al.
Published: (2023)
by: Piratla, Vihari, et al.
Published: (2023)
Evaluating the Relevance of Uncertainty Estimators for LLM Hallucination
by: Agnimo, Yedidia, et al.
Published: (2026)
by: Agnimo, Yedidia, et al.
Published: (2026)
Dual-Uncertainty Guided Policy Learning for Multimodal Reasoning
by: Liu, Rui, et al.
Published: (2025)
by: Liu, Rui, et al.
Published: (2025)
Capability-Based Scaling Trends for LLM-Based Red-Teaming
by: Panfilov, Alexander, et al.
Published: (2025)
by: Panfilov, Alexander, et al.
Published: (2025)
BayesAgent: Bayesian Agentic Reasoning Under Uncertainty via Verbalized Probabilistic Graphical Modeling
by: Huang, Hengguan, et al.
Published: (2024)
by: Huang, Hengguan, et al.
Published: (2024)
Scaling Reasoning without Attention
by: Zhao, Xueliang, et al.
Published: (2025)
by: Zhao, Xueliang, et al.
Published: (2025)
MathScale: Scaling Instruction Tuning for Mathematical Reasoning
by: Tang, Zhengyang, et al.
Published: (2024)
by: Tang, Zhengyang, et al.
Published: (2024)
Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty
by: Damani, Mehul, et al.
Published: (2025)
by: Damani, Mehul, et al.
Published: (2025)
Dist2ill: Distributional Distillation for One-Pass Uncertainty Estimation in Large Language Models
by: Zhao, Yicong, et al.
Published: (2025)
by: Zhao, Yicong, et al.
Published: (2025)
Low-Rank Quantization-Aware Training for LLMs
by: Bondarenko, Yelysei, et al.
Published: (2024)
by: Bondarenko, Yelysei, et al.
Published: (2024)
Predicting Satisfaction of Counterfactual Explanations from Human Ratings of Explanatory Qualities
by: Domnich, Marharyta, et al.
Published: (2025)
by: Domnich, Marharyta, et al.
Published: (2025)
ScaleDiff: Scaling Difficult Problems for Advanced Mathematical Reasoning
by: Pei, Qizhi, et al.
Published: (2025)
by: Pei, Qizhi, et al.
Published: (2025)
Recitation over Reasoning: How Cutting-Edge Language Models Can Fail on Elementary School-Level Reasoning Problems?
by: Yan, Kai, et al.
Published: (2025)
by: Yan, Kai, et al.
Published: (2025)
Group Reasoning Emission Estimation Networks
by: Guo, Yanming, et al.
Published: (2025)
by: Guo, Yanming, et al.
Published: (2025)
Test-Time Scaling in Reasoning Models Is Not Effective for Knowledge-Intensive Tasks Yet
by: Zhao, James Xu, et al.
Published: (2025)
by: Zhao, James Xu, et al.
Published: (2025)
Nemotron-Cascade: Scaling Cascaded Reinforcement Learning for General-Purpose Reasoning Models
by: Wang, Boxin, et al.
Published: (2025)
by: Wang, Boxin, et al.
Published: (2025)
On the Role of Temperature Sampling in Test-Time Scaling
by: Wu, Yuheng, et al.
Published: (2025)
by: Wu, Yuheng, et al.
Published: (2025)
Similar Items
-
On the Usefulness of the Fit-on-the-Test View on Evaluating Calibration of Classifiers
by: Kängsepp, Markus, et al.
Published: (2022) -
Enhancing Counterfactual Explanation Search with Diffusion Distance and Directional Coherence
by: Domnich, Marharyta, et al.
Published: (2024) -
Cross-model Fairness: Empirical Study of Fairness and Ethics Under Model Multiplicity
by: Sokol, Kacper, et al.
Published: (2022) -
Cautious Calibration in Binary Classification
by: Allikivi, Mari-Liis, et al.
Published: (2024) -
Towards Unifying Evaluation of Counterfactual Explanations: Leveraging Large Language Models for Human-Centric Assessments
by: Domnich, Marharyta, et al.
Published: (2024)