Certain but not Probable? Differentiating Certainty from Probability in LLM Token Outputs for Probabilistic Scenarios
Fuente:
arXiv
Saved in:
| Main Authors: | Toney-Wails, Autumn, Wails, Ryan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
How Clued up are LLMs? Evaluating Multi-Step Deductive Reasoning in a Text-Based Game Environment
by: Ansell, Rebecca, et al.
Published: (2026)
by: Ansell, Rebecca, et al.
Published: (2026)
AI on AI: Exploring the Utility of GPT as an Expert Annotator of AI Publications
by: Toney-Wails, Autumn, et al.
Published: (2024)
by: Toney-Wails, Autumn, et al.
Published: (2024)
Understanding Token Probability Encoding in Output Embeddings
by: Cho, Hakaze, et al.
Published: (2024)
by: Cho, Hakaze, et al.
Published: (2024)
A Measurement of Genuine Tor Traces for Realistic Website Fingerprinting
by: Jansen, Rob, et al.
Published: (2024)
by: Jansen, Rob, et al.
Published: (2024)
Reducing the Probability of Undesirable Outputs in Language Models Using Probabilistic Inference
by: Zhao, Stephen, et al.
Published: (2025)
by: Zhao, Stephen, et al.
Published: (2025)
Training-free LLM-generated Text Detection by Mining Token Probability Sequences
by: Xu, Yihuai, et al.
Published: (2024)
by: Xu, Yihuai, et al.
Published: (2024)
Probability Distributions Computed by Autoregressive Transformers
by: Yang, Andy, et al.
Published: (2025)
by: Yang, Andy, et al.
Published: (2025)
Probability of Differentiation Reveals Brittleness of Homogeneity Bias in GPT-4
by: Lee, Messi H. J., et al.
Published: (2024)
by: Lee, Messi H. J., et al.
Published: (2024)
First Token Probability Guided RAG for Telecom Question Answering
by: Chen, Tingwei, et al.
Published: (2025)
by: Chen, Tingwei, et al.
Published: (2025)
TPA: Next Token Probability Attribution for Detecting Hallucinations in RAG
by: Lu, Pengqian, et al.
Published: (2025)
by: Lu, Pengqian, et al.
Published: (2025)
Probability-Consistent Preference Optimization for Enhanced LLM Reasoning
by: Yang, Yunqiao, et al.
Published: (2025)
by: Yang, Yunqiao, et al.
Published: (2025)
On the Salience of Low-Probability Tokens for AI-Generated Text Detection: A Multiscale Uncertainty Perspective
by: Guo, Yikai, et al.
Published: (2026)
by: Guo, Yikai, et al.
Published: (2026)
Fact-Checking with Large Language Models via Probabilistic Certainty and Consistency
by: Wang, Haoran, et al.
Published: (2026)
by: Wang, Haoran, et al.
Published: (2026)
Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs
by: Yang, Zhihe, et al.
Published: (2025)
by: Yang, Zhihe, et al.
Published: (2025)
Exact Byte-Level Probabilities from Tokenized Language Models for FIM-Tasks and Model Ensembles
by: Phan, Buu, et al.
Published: (2024)
by: Phan, Buu, et al.
Published: (2024)
Attention2Probability: Attention-Driven Terminology Probability Estimation for Robust Speech-to-Text System
by: Du, Yanfan, et al.
Published: (2025)
by: Du, Yanfan, et al.
Published: (2025)
Alignment-Enhanced Decoding:Defending via Token-Level Adaptive Refining of Probability Distributions
by: Liu, Quan, et al.
Published: (2024)
by: Liu, Quan, et al.
Published: (2024)
Textual Entailment is not a Better Bias Metric than Token Probability
by: Felkner, Virginia K., et al.
Published: (2025)
by: Felkner, Virginia K., et al.
Published: (2025)
How to Compute the Probability of a Word
by: Pimentel, Tiago, et al.
Published: (2024)
by: Pimentel, Tiago, et al.
Published: (2024)
"My Answer is C": First-Token Probabilities Do Not Match Text Answers in Instruction-Tuned Language Models
by: Wang, Xinpeng, et al.
Published: (2024)
by: Wang, Xinpeng, et al.
Published: (2024)
ProFit: Leveraging High-Value Signals in SFT via Probability-Guided Token Selection
by: Liu, Tao, et al.
Published: (2026)
by: Liu, Tao, et al.
Published: (2026)
Zonkey: A Hierarchical Diffusion Language Model with Differentiable Tokenization and Probabilistic Attention
by: Rozental, Alon
Published: (2026)
by: Rozental, Alon
Published: (2026)
LLM Probability Concentration: How Alignment Shrinks the Generative Horizon
by: Yang, Chenghao, et al.
Published: (2025)
by: Yang, Chenghao, et al.
Published: (2025)
Rehearsing Answers to Probable Questions with Perspective-Taking
by: Shih, Yung-Yu, et al.
Published: (2024)
by: Shih, Yung-Yu, et al.
Published: (2024)
Position Paper On Diagnostic Uncertainty Estimation from Large Language Models: Next-Word Probability Is Not Pre-test Probability
by: Gao, Yanjun, et al.
Published: (2024)
by: Gao, Yanjun, et al.
Published: (2024)
A Probability--Quality Trade-off in Aligned Language Models and its Relation to Sampling Adaptors
by: Tan, Naaman, et al.
Published: (2024)
by: Tan, Naaman, et al.
Published: (2024)
Bridging Internal Probability and Self-Consistency for Effective and Efficient LLM Reasoning
by: Zhou, Zhi, et al.
Published: (2025)
by: Zhou, Zhi, et al.
Published: (2025)
Detecting Hallucinations in Large Language Model Generation: A Token Probability Approach
by: Quevedo, Ernesto, et al.
Published: (2024)
by: Quevedo, Ernesto, et al.
Published: (2024)
Safeguarding RAG Pipelines with GMTP: A Gradient-based Masked Token Probability Method for Poisoned Document Detection
by: Kim, San, et al.
Published: (2025)
by: Kim, San, et al.
Published: (2025)
Calibrating Expressions of Certainty
by: Wang, Peiqi, et al.
Published: (2024)
by: Wang, Peiqi, et al.
Published: (2024)
How Should We Model the Probability of a Language?
by: Dent, Rasul, et al.
Published: (2026)
by: Dent, Rasul, et al.
Published: (2026)
Back to Basics: Revisiting Exploration in Reinforcement Learning for LLM Reasoning via Generative Probabilities
by: Li, Pengyi, et al.
Published: (2026)
by: Li, Pengyi, et al.
Published: (2026)
Improving LLM First-Token Predictions in Multiple-Choice Question Answering via Output Prefilling
by: Cappelletti, Silvia, et al.
Published: (2025)
by: Cappelletti, Silvia, et al.
Published: (2025)
Certainty robustness: Evaluating LLM stability under self-challenging prompts
by: Saadat, Mohammadreza, et al.
Published: (2026)
by: Saadat, Mohammadreza, et al.
Published: (2026)
Calibrating Verbalized Probabilities for Large Language Models
by: Wang, Cheng, et al.
Published: (2024)
by: Wang, Cheng, et al.
Published: (2024)
Incoherent Probability Judgments in Large Language Models
by: Zhu, Jian-Qiao, et al.
Published: (2024)
by: Zhu, Jian-Qiao, et al.
Published: (2024)
PRISM: Probability Reallocation with In-Span Masking for Knowledge-Sensitive Alignment
by: Xu, Chenning, et al.
Published: (2026)
by: Xu, Chenning, et al.
Published: (2026)
Finding Replicable Human Evaluations via Stable Ranking Probability
by: Riley, Parker, et al.
Published: (2024)
by: Riley, Parker, et al.
Published: (2024)
Beyond Probabilities: Unveiling the Misalignment in Evaluating Large Language Models
by: Lyu, Chenyang, et al.
Published: (2024)
by: Lyu, Chenyang, et al.
Published: (2024)
Interpreting Predictive Probabilities: Model Confidence or Human Label Variation?
by: Baan, Joris, et al.
Published: (2024)
by: Baan, Joris, et al.
Published: (2024)
Similar Items
-
How Clued up are LLMs? Evaluating Multi-Step Deductive Reasoning in a Text-Based Game Environment
by: Ansell, Rebecca, et al.
Published: (2026) -
AI on AI: Exploring the Utility of GPT as an Expert Annotator of AI Publications
by: Toney-Wails, Autumn, et al.
Published: (2024) -
Understanding Token Probability Encoding in Output Embeddings
by: Cho, Hakaze, et al.
Published: (2024) -
A Measurement of Genuine Tor Traces for Realistic Website Fingerprinting
by: Jansen, Rob, et al.
Published: (2024) -
Reducing the Probability of Undesirable Outputs in Language Models Using Probabilistic Inference
by: Zhao, Stephen, et al.
Published: (2025)