Entropy trajectory shape predicts LLM reasoning reliability: A diagnostic study of uncertainty dynamics in chain-of-thought
Fuente:
arXiv
Enregistré dans:
| Auteur principal: | Zhao, Xinghao |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
How to think step-by-step: A mechanistic understanding of chain-of-thought reasoning
par: Dutta, Subhabrata, et autres
Publié: (2024)
par: Dutta, Subhabrata, et autres
Publié: (2024)
How well can off-the-shelf LLMs elucidate molecular structures from mass spectra using chain-of-thought reasoning?
par: Wang, Yufeng, et autres
Publié: (2026)
par: Wang, Yufeng, et autres
Publié: (2026)
To CoT or not to CoT? Chain-of-thought helps mainly on math and symbolic reasoning
par: Sprague, Zayne, et autres
Publié: (2024)
par: Sprague, Zayne, et autres
Publié: (2024)
Your thoughts tell who you are: Characterize the reasoning patterns of LRMs
par: Chen, Yida, et autres
Publié: (2025)
par: Chen, Yida, et autres
Publié: (2025)
Encode, Think, Decode: Scaling test-time reasoning with recursive latent thoughts
par: Koishekenov, Yeskendir, et autres
Publié: (2025)
par: Koishekenov, Yeskendir, et autres
Publié: (2025)
Measuring and curing reasoning rigidity: from decorative chain-of-thought to genuine faithfulness
par: Basu, Abhinaba, et autres
Publié: (2026)
par: Basu, Abhinaba, et autres
Publié: (2026)
Large language models can learn and generalize steganographic chain-of-thought under process supervision
par: Skaf, Joey, et autres
Publié: (2025)
par: Skaf, Joey, et autres
Publié: (2025)
Activation Control for Efficiently Eliciting Long Chain-of-thought Ability of Language Models
par: Zhao, Zekai, et autres
Publié: (2025)
par: Zhao, Zekai, et autres
Publié: (2025)
Probabilistic unifying relations for modelling epistemic and aleatoric uncertainty: semantics and automated reasoning with theorem proving
par: Ye, Kangfeng, et autres
Publié: (2023)
par: Ye, Kangfeng, et autres
Publié: (2023)
EPO: Entropy-regularized Policy Optimization for LLM Agents Reinforcement Learning
par: Xu, Wujiang, et autres
Publié: (2025)
par: Xu, Wujiang, et autres
Publié: (2025)
AI-generated data contamination erodes pathological variability and diagnostic reliability
par: He, Hongyu, et autres
Publié: (2026)
par: He, Hongyu, et autres
Publié: (2026)
OCEAN: Offline Chain-of-thought Evaluation and Alignment in Large Language Models
par: Wu, Junda, et autres
Publié: (2024)
par: Wu, Junda, et autres
Publié: (2024)
When an LLM is apprehensive about its answers -- and when its uncertainty is justified
par: Sychev, Petr, et autres
Publié: (2025)
par: Sychev, Petr, et autres
Publié: (2025)
Entropy Law: The Story Behind Data Compression and LLM Performance
par: Yin, Mingjia, et autres
Publié: (2024)
par: Yin, Mingjia, et autres
Publié: (2024)
On multi-token prediction for efficient LLM inference
par: Mehra, Somesh, et autres
Publié: (2025)
par: Mehra, Somesh, et autres
Publié: (2025)
Can Language Models Perform Robust Reasoning in Chain-of-thought Prompting with Noisy Rationales?
par: Zhou, Zhanke, et autres
Publié: (2024)
par: Zhou, Zhanke, et autres
Publié: (2024)
Sudoku-Bench: Evaluating creative reasoning with Sudoku variants
par: Seely, Jeffrey, et autres
Publié: (2025)
par: Seely, Jeffrey, et autres
Publié: (2025)
Addressing Performance Saturation for LLM RL via Precise Entropy Curve Control
par: Li, Bolian, et autres
Publié: (2026)
par: Li, Bolian, et autres
Publié: (2026)
Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents
par: Wang, Jiawei, et autres
Publié: (2025)
par: Wang, Jiawei, et autres
Publié: (2025)
Beyond Semantic Entropy: Boosting LLM Uncertainty Quantification with Pairwise Semantic Similarity
par: Nguyen, Dang, et autres
Publié: (2025)
par: Nguyen, Dang, et autres
Publié: (2025)
Language models scale reliably with over-training and on downstream tasks
par: Gadre, Samir Yitzhak, et autres
Publié: (2024)
par: Gadre, Samir Yitzhak, et autres
Publié: (2024)
Boosting classification reliability of NLP transformer models in the long run
par: Kmetty, Zoltán, et autres
Publié: (2023)
par: Kmetty, Zoltán, et autres
Publié: (2023)
Intrinsic Entropy of Context Length Scaling in LLMs
par: Shi, Jingzhe, et autres
Publié: (2025)
par: Shi, Jingzhe, et autres
Publié: (2025)
Language hooks: a modular framework for augmenting LLM reasoning that decouples tool usage from the model and its prompt
par: de Mijolla, Damien, et autres
Publié: (2024)
par: de Mijolla, Damien, et autres
Publié: (2024)
When can transformers reason with abstract symbols?
par: Boix-Adsera, Enric, et autres
Publié: (2023)
par: Boix-Adsera, Enric, et autres
Publié: (2023)
Artificial Expert Intelligence through PAC-reasoning
par: Shalev-Shwartz, Shai, et autres
Publié: (2024)
par: Shalev-Shwartz, Shai, et autres
Publié: (2024)
Navigation under uncertainty: Trajectory prediction and occlusion reasoning with switching dynamical systems
par: Wei, Ran, et autres
Publié: (2024)
par: Wei, Ran, et autres
Publié: (2024)
Density estimation with LLMs: a geometric investigation of in-context learning trajectories
par: Liu, Toni J. B., et autres
Publié: (2024)
par: Liu, Toni J. B., et autres
Publié: (2024)
Is continuous CoT better suited for multi-lingual reasoning?
par: Bashir, Ali Hamza, et autres
Publié: (2026)
par: Bashir, Ali Hamza, et autres
Publié: (2026)
Are complicated loss functions necessary for teaching LLMs to reason?
par: Carrino, Gabriele, et autres
Publié: (2026)
par: Carrino, Gabriele, et autres
Publié: (2026)
Statistical multi-metric evaluation and visualization of LLM system predictive performance
par: Ackerman, Samuel, et autres
Publié: (2025)
par: Ackerman, Samuel, et autres
Publié: (2025)
Few shot chain-of-thought driven reasoning to prompt LLMs for open ended medical question answering
par: Nachane, Saeel Sandeep, et autres
Publié: (2024)
par: Nachane, Saeel Sandeep, et autres
Publié: (2024)
AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning
par: Wang, Tevin, et autres
Publié: (2025)
par: Wang, Tevin, et autres
Publié: (2025)
Neural networks for abstraction and reasoning: Towards broad generalization in machines
par: Bober-Irizar, Mikel, et autres
Publié: (2024)
par: Bober-Irizar, Mikel, et autres
Publié: (2024)
Functional Entropy: Predicting Functional Correctness in LLM-Generated Code with Uncertainty Quantification
par: Bouchard, Dylan, et autres
Publié: (2026)
par: Bouchard, Dylan, et autres
Publié: (2026)
Explore Briefly, Then Decide: Mitigating LLM Overthinking via Cumulative Entropy Regulation
par: Bin, Yi, et autres
Publié: (2025)
par: Bin, Yi, et autres
Publié: (2025)
LANCET: Neural Intervention via Structural Entropy for Mitigating Faithfulness Hallucinations in LLMs
par: Wang, Chenxu, et autres
Publié: (2026)
par: Wang, Chenxu, et autres
Publié: (2026)
Entropy-Regularized Process Reward Model
par: Zhang, Hanning, et autres
Publié: (2024)
par: Zhang, Hanning, et autres
Publié: (2024)
Failure Modes of Maximum Entropy RLHF
par: Çağatan, Ömer Veysel, et autres
Publié: (2025)
par: Çağatan, Ömer Veysel, et autres
Publié: (2025)
Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control
par: Zhang, Jiazheng, et autres
Publié: (2026)
par: Zhang, Jiazheng, et autres
Publié: (2026)
Documents similaires
-
How to think step-by-step: A mechanistic understanding of chain-of-thought reasoning
par: Dutta, Subhabrata, et autres
Publié: (2024) -
How well can off-the-shelf LLMs elucidate molecular structures from mass spectra using chain-of-thought reasoning?
par: Wang, Yufeng, et autres
Publié: (2026) -
To CoT or not to CoT? Chain-of-thought helps mainly on math and symbolic reasoning
par: Sprague, Zayne, et autres
Publié: (2024) -
Your thoughts tell who you are: Characterize the reasoning patterns of LRMs
par: Chen, Yida, et autres
Publié: (2025) -
Encode, Think, Decode: Scaling test-time reasoning with recursive latent thoughts
par: Koishekenov, Yeskendir, et autres
Publié: (2025)