Confidence Improves Self-Consistency in LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Taubenfeld, Amir, Sheffer, Tom, Ofek, Eran, Feder, Amir, Goldstein, Ariel, Gekhman, Zorik, Yona, Gal |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Empty Shelves or Lost Keys? Recall Is the Bottleneck for Parametric Factuality
by: Calderon, Nitay, et al.
Published: (2026)
by: Calderon, Nitay, et al.
Published: (2026)
Can LLMs Learn Macroeconomic Narratives from Social Media?
by: Gueta, Almog, et al.
Published: (2024)
by: Gueta, Almog, et al.
Published: (2024)
Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?
by: Gekhman, Zorik, et al.
Published: (2024)
by: Gekhman, Zorik, et al.
Published: (2024)
Systematic Biases in LLM Simulations of Debates
by: Taubenfeld, Amir, et al.
Published: (2024)
by: Taubenfeld, Amir, et al.
Published: (2024)
Evaluating Alignment of Behavioral Dispositions in LLMs
by: Taubenfeld, Amir, et al.
Published: (2026)
by: Taubenfeld, Amir, et al.
Published: (2026)
Thinking to Recall: How Reasoning Unlocks Parametric Knowledge in LLMs
by: Gekhman, Zorik, et al.
Published: (2026)
by: Gekhman, Zorik, et al.
Published: (2026)
Do LLMs have Consistent Values?
by: Rozen, Naama, et al.
Published: (2024)
by: Rozen, Naama, et al.
Published: (2024)
Distributional reasoning in LLMs: Parallel reasoning processes in multi-hop reasoning
by: Shalev, Yuval, et al.
Published: (2024)
by: Shalev, Yuval, et al.
Published: (2024)
Inside-Out: Hidden Factual Knowledge in LLMs
by: Gekhman, Zorik, et al.
Published: (2025)
by: Gekhman, Zorik, et al.
Published: (2025)
Exploring the Learning Capabilities of Language Models using LEVERWORLDS
by: Wagner, Eitan, et al.
Published: (2024)
by: Wagner, Eitan, et al.
Published: (2024)
LLMs Know More Than They Show: On the Intrinsic Representation of LLM Hallucinations
by: Orgad, Hadas, et al.
Published: (2024)
by: Orgad, Hadas, et al.
Published: (2024)
NL-Eye: Abductive NLI for Images
by: Ventura, Mor, et al.
Published: (2024)
by: Ventura, Mor, et al.
Published: (2024)
Keep Guessing? When Considering Inference Scaling, Mind the Baselines
by: Yona, Gal, et al.
Published: (2024)
by: Yona, Gal, et al.
Published: (2024)
Improving the Reliability of LLMs: Combining CoT, RAG, Self-Consistency, and Self-Verification
by: Kumar, Adarsh, et al.
Published: (2025)
by: Kumar, Adarsh, et al.
Published: (2025)
Why Fine-Tuning Encourages Hallucinations and How to Fix It
by: Kaplan, Guy, et al.
Published: (2026)
by: Kaplan, Guy, et al.
Published: (2026)
In-Context Representation Hijacking
by: Yona, Itay, et al.
Published: (2025)
by: Yona, Itay, et al.
Published: (2025)
SAUCE: Synchronous and Asynchronous User-Customizable Environment for Multi-Agent LLM Interaction
by: Neuberger, Shlomo, et al.
Published: (2024)
by: Neuberger, Shlomo, et al.
Published: (2024)
Self-Training Meets Consistency: Improving LLMs' Reasoning with Consistency-Driven Rationale Evaluation
by: Lee, Jaehyeok, et al.
Published: (2024)
by: Lee, Jaehyeok, et al.
Published: (2024)
Advancing NLP Security by Leveraging LLMs as Adversarial Engines
by: Srinivasan, Sudarshan, et al.
Published: (2024)
by: Srinivasan, Sudarshan, et al.
Published: (2024)
Fine-Grained Detection of Context-Grounded Hallucinations Using LLMs
by: Peisakhovsky, Yehonatan, et al.
Published: (2025)
by: Peisakhovsky, Yehonatan, et al.
Published: (2025)
Evaluating Prompt Engineering Techniques for Accuracy and Confidence Elicitation in Medical LLMs
by: Naderi, Nariman, et al.
Published: (2025)
by: Naderi, Nariman, et al.
Published: (2025)
CoRefine: Confidence-Guided Self-Refinement for Adaptive Test-Time Compute
by: Jin, Chen, et al.
Published: (2026)
by: Jin, Chen, et al.
Published: (2026)
Do Language Models Need Sleep? Offline Recurrence for Improved Online Inference
by: Lee, Sangyun, et al.
Published: (2026)
by: Lee, Sangyun, et al.
Published: (2026)
The Self-Execution Benchmark: Measuring LLMs' Attempts to Overcome Their Lack of Self-Execution
by: Ezra, Elon, et al.
Published: (2025)
by: Ezra, Elon, et al.
Published: (2025)
Redefining "Hallucination" in LLMs: Towards a psychology-informed framework for mitigating misinformation
by: Berberette, Elijah, et al.
Published: (2024)
by: Berberette, Elijah, et al.
Published: (2024)
DRAGged into Conflicts: Detecting and Addressing Conflicting Sources in Search-Augmented LLMs
by: Cattan, Arie, et al.
Published: (2025)
by: Cattan, Arie, et al.
Published: (2025)
Beyond Self-Consistency: Ensemble Reasoning Boosts Consistency and Accuracy of LLMs in Cancer Staging
by: Chang, Chia-Hsuan, et al.
Published: (2024)
by: Chang, Chia-Hsuan, et al.
Published: (2024)
SaySelf: Teaching LLMs to Express Confidence with Self-Reflective Rationales
by: Xu, Tianyang, et al.
Published: (2024)
by: Xu, Tianyang, et al.
Published: (2024)
From Building Blocks to Planning: Multi-Step Spatial Reasoning in LLMs with Reinforcement Learning
by: Tahmasbi, Amir, et al.
Published: (2025)
by: Tahmasbi, Amir, et al.
Published: (2025)
Beyond Confidence: Rethinking Self-Assessments for Performance Prediction in LLMs
by: Bhattacharyya, Sree, et al.
Published: (2026)
by: Bhattacharyya, Sree, et al.
Published: (2026)
Improving Multi-turn Dialogue Consistency with Self-Recall Thinking
by: Pang, Renning, et al.
Published: (2026)
by: Pang, Renning, et al.
Published: (2026)
Multi-Perspective Consistency Enhances Confidence Estimation in Large Language Models
by: Wang, Pei, et al.
Published: (2024)
by: Wang, Pei, et al.
Published: (2024)
Majority of the Bests: Improving Best-of-N via Bootstrapping
by: Rakhsha, Amin, et al.
Published: (2025)
by: Rakhsha, Amin, et al.
Published: (2025)
A Dual-Axis Taxonomy of Knowledge Editing for LLMs: From Mechanisms to Functions
by: Salehoof, Amir Mohammad, et al.
Published: (2025)
by: Salehoof, Amir Mohammad, et al.
Published: (2025)
Contradiction Detection in RAG Systems: Evaluating LLMs as Context Validators for Improved Information Consistency
by: Gokul, Vignesh, et al.
Published: (2025)
by: Gokul, Vignesh, et al.
Published: (2025)
Reversing Large Language Models for Efficient Training and Fine-Tuning
by: Gal, Eshed, et al.
Published: (2025)
by: Gal, Eshed, et al.
Published: (2025)
Bias Similarity Measurement: A Black-Box Audit of Fairness Across LLMs
by: Jeong, Hyejun, et al.
Published: (2024)
by: Jeong, Hyejun, et al.
Published: (2024)
Persian-Phi: Efficient Cross-Lingual Adaptation of Compact LLMs via Curriculum Learning
by: Akhlaghi, Amir Mohammad, et al.
Published: (2025)
by: Akhlaghi, Amir Mohammad, et al.
Published: (2025)
Confident RAG: Enhancing the Performance of LLMs for Mathematics Question Answering through Multi-Embedding and Confidence Scoring
by: Chen, Shiting, et al.
Published: (2025)
by: Chen, Shiting, et al.
Published: (2025)
Self-Consistency Is Losing Its Edge: Diminishing Returns and Rising Costs in Modern LLMs
by: Loo, Chiyan
Published: (2025)
by: Loo, Chiyan
Published: (2025)
Similar Items
-
Empty Shelves or Lost Keys? Recall Is the Bottleneck for Parametric Factuality
by: Calderon, Nitay, et al.
Published: (2026) -
Can LLMs Learn Macroeconomic Narratives from Social Media?
by: Gueta, Almog, et al.
Published: (2024) -
Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?
by: Gekhman, Zorik, et al.
Published: (2024) -
Systematic Biases in LLM Simulations of Debates
by: Taubenfeld, Amir, et al.
Published: (2024) -
Evaluating Alignment of Behavioral Dispositions in LLMs
by: Taubenfeld, Amir, et al.
Published: (2026)