Zero-Shot Confidence Estimation for Small LLMs: When Supervised Baselines Aren't Worth Training
Fuente:
arXiv
Saved in:
| Main Author: | Nguyen, Luong N. |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
HalluShift: Measuring Distribution Shifts towards Hallucination Detection in LLMs
by: Dasgupta, Sharanya, et al.
Published: (2025)
by: Dasgupta, Sharanya, et al.
Published: (2025)
G-Zero: Self-Play for Open-Ended Generation from Zero Data
by: Huang, Chengsong, et al.
Published: (2026)
by: Huang, Chengsong, et al.
Published: (2026)
When to Reason: Semantic Router for vLLM
by: Wang, Chen, et al.
Published: (2025)
by: Wang, Chen, et al.
Published: (2025)
Bayesian Orchestration of Multi-LLM Agents for Cost-Aware Sequential Decision-Making
by: Amin, Danial
Published: (2026)
by: Amin, Danial
Published: (2026)
Multilingual Machine Translation with Quantum Encoder Decoder Attention-based Convolutional Variational Circuits
by: Dikshit, Subrit, et al.
Published: (2025)
by: Dikshit, Subrit, et al.
Published: (2025)
The Future of MLLM Prompting is Adaptive: A Comprehensive Experimental Evaluation of Prompt Engineering Methods for Robust Multimodal Performance
by: Mohanty, Anwesha, et al.
Published: (2025)
by: Mohanty, Anwesha, et al.
Published: (2025)
Are Large Language Models Reliable Argument Quality Annotators?
by: Mirzakhmedova, Nailia, et al.
Published: (2024)
by: Mirzakhmedova, Nailia, et al.
Published: (2024)
Iterative Prompting with Persuasion Skills in Jailbreaking Large Language Models
by: Ke, Shih-Wen, et al.
Published: (2025)
by: Ke, Shih-Wen, et al.
Published: (2025)
A Novel Nuanced Conversation Evaluation Framework for Large Language Models in Mental Health
by: Marrapese, Alexander, et al.
Published: (2024)
by: Marrapese, Alexander, et al.
Published: (2024)
How Many Bytes Can You Take Out Of Brain-To-Text Decoding?
by: Antonello, Richard, et al.
Published: (2024)
by: Antonello, Richard, et al.
Published: (2024)
Artificial Agency and Large Language Models
by: van Lier, Maud, et al.
Published: (2024)
by: van Lier, Maud, et al.
Published: (2024)
Can Large Language Models Act as Symbolic Reasoners?
by: Sullivan, Rob, et al.
Published: (2024)
by: Sullivan, Rob, et al.
Published: (2024)
Quantifying the Effectiveness of Student Organization Activities using Natural Language Processing
by: Taruc, Lyberius Ennio F., et al.
Published: (2024)
by: Taruc, Lyberius Ennio F., et al.
Published: (2024)
Neurosymbolic Graph Enrichment for Grounded World Models
by: De Giorgis, Stefano, et al.
Published: (2024)
by: De Giorgis, Stefano, et al.
Published: (2024)
Training-Free Agentic AI: Probabilistic Control and Coordination in Multi-Agent LLM Systems
by: Hosseini, Mohammad Parsa, et al.
Published: (2026)
by: Hosseini, Mohammad Parsa, et al.
Published: (2026)
Automated Thematic Analyses Using LLMs: Xylazine Wound Management Social Media Chatter Use Case
by: Hairston, JaMor, et al.
Published: (2025)
by: Hairston, JaMor, et al.
Published: (2025)
Multilingual Text-to-SQL: Benchmarking the Limits of Language Models with Collaborative Language Agents
by: Pham, Khanh Trinh, et al.
Published: (2025)
by: Pham, Khanh Trinh, et al.
Published: (2025)
Language-Dependent Political Bias in AI: A Study of ChatGPT and Gemini
by: Yuksel, Dogus, et al.
Published: (2025)
by: Yuksel, Dogus, et al.
Published: (2025)
Improving Object Detector Training on Synthetic Data by Starting With a Strong Baseline Methodology
by: Ruis, Frank A., et al.
Published: (2024)
by: Ruis, Frank A., et al.
Published: (2024)
Use of AI Tools: Guidelines to Maintain Academic Integrity in Computing Colleges
by: El-boghdadi, Hatem M., et al.
Published: (2026)
by: El-boghdadi, Hatem M., et al.
Published: (2026)
A Review of Challenges in Speech-based Conversational AI for Elderly Care
by: Klaassen, Willemijn, et al.
Published: (2024)
by: Klaassen, Willemijn, et al.
Published: (2024)
Can Large Language Models Understand As Well As Apply Patent Regulations to Pass a Hands-On Patent Attorney Test?
by: Khera, Bhakti, et al.
Published: (2025)
by: Khera, Bhakti, et al.
Published: (2025)
Reflexive Prompt Engineering: A Framework for Responsible Prompt Engineering and Interaction Design
by: Djeffal, Christian
Published: (2025)
by: Djeffal, Christian
Published: (2025)
Reasoning Abilities of Large Language Models: In-Depth Analysis on the Abstraction and Reasoning Corpus
by: Lee, Seungpil, et al.
Published: (2024)
by: Lee, Seungpil, et al.
Published: (2024)
Which English Do LLMs Prefer? Triangulating Structural Bias Towards American English in Foundation Models
by: Nayeem, Mir Tafseer, et al.
Published: (2026)
by: Nayeem, Mir Tafseer, et al.
Published: (2026)
Confidence Interval Estimation of Predictive Performance in the Context of AutoML
by: Paraschakis, Konstantinos, et al.
Published: (2024)
by: Paraschakis, Konstantinos, et al.
Published: (2024)
ClinEnv: An Interactive Multi-Stage Long Horizon EHR Environment for Agents
by: Lu, Yuxing, et al.
Published: (2026)
by: Lu, Yuxing, et al.
Published: (2026)
"Do Not Mention This to the User": Detecting and Understanding Malicious Agent Skills
by: Liu, Yi, et al.
Published: (2026)
by: Liu, Yi, et al.
Published: (2026)
Scaling Multiagent Systems with Process Rewards
by: Li, Ed, et al.
Published: (2026)
by: Li, Ed, et al.
Published: (2026)
AgentRedBench: Dynamic Redteaming and Integration-Aware Defense for LLM Agents over SaaS Integrations
by: Dingeto, Hiskias, et al.
Published: (2026)
by: Dingeto, Hiskias, et al.
Published: (2026)
InvestAlign: Overcoming Data Scarcity in Aligning Large Language Models with Investor Decision-Making Processes under Herd Behavior
by: Wang, Huisheng, et al.
Published: (2025)
by: Wang, Huisheng, et al.
Published: (2025)
On the Limitations of Compute Thresholds as a Governance Strategy
by: Hooker, Sara
Published: (2024)
by: Hooker, Sara
Published: (2024)
An Autonomous GIS Agent Framework for Geospatial Data Retrieval
by: Ning, Huan, et al.
Published: (2024)
by: Ning, Huan, et al.
Published: (2024)
Teaching Programming in the Age of Generative AI: Insights from Literature, Pedagogical Proposals, and Student Perspectives
by: Rubio-Manzano, Clemente, et al.
Published: (2025)
by: Rubio-Manzano, Clemente, et al.
Published: (2025)
Effects of Prompt Length on Domain-specific Tasks for Large Language Models
by: Liu, Qibang, et al.
Published: (2025)
by: Liu, Qibang, et al.
Published: (2025)
Automatic Generation of Behavioral Test Cases For Natural Language Processing Using Clustering and Prompting
by: Li, Ying, et al.
Published: (2024)
by: Li, Ying, et al.
Published: (2024)
Combinatorial Reasoning: Selecting Reasons in Generative AI Pipelines via Combinatorial Optimization
by: Esencan, Mert, et al.
Published: (2024)
by: Esencan, Mert, et al.
Published: (2024)
Is Model Collapse Inevitable? Breaking the Curse of Recursion by Accumulating Real and Synthetic Data
by: Gerstgrasser, Matthias, et al.
Published: (2024)
by: Gerstgrasser, Matthias, et al.
Published: (2024)
An Auditable Pipeline for Fuzzy Full-Text Screening in Systematic Reviews: Integrating Contrastive Semantic Highlighting and LLM Judgment
by: Mortezaagha, Pouria, et al.
Published: (2025)
by: Mortezaagha, Pouria, et al.
Published: (2025)
Embedding-Aligned Language Models
by: Tennenholtz, Guy, et al.
Published: (2024)
by: Tennenholtz, Guy, et al.
Published: (2024)
Similar Items
-
HalluShift: Measuring Distribution Shifts towards Hallucination Detection in LLMs
by: Dasgupta, Sharanya, et al.
Published: (2025) -
G-Zero: Self-Play for Open-Ended Generation from Zero Data
by: Huang, Chengsong, et al.
Published: (2026) -
When to Reason: Semantic Router for vLLM
by: Wang, Chen, et al.
Published: (2025) -
Bayesian Orchestration of Multi-LLM Agents for Cost-Aware Sequential Decision-Making
by: Amin, Danial
Published: (2026) -
Multilingual Machine Translation with Quantum Encoder Decoder Attention-based Convolutional Variational Circuits
by: Dikshit, Subrit, et al.
Published: (2025)