Calibrating LLM Confidence by Probing Perturbed Representation Stability
Fuente:
arXiv
Saved in:
| Main Authors: | Khanmohammadi, Reza, Miahi, Erfan, Mardikoraem, Mehrsa, Kaur, Simerjot, Brugere, Ivan, Smiley, Charese H., Thind, Kundan, Ghassemi, Mohammad M. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Grounded or Guessing? LVLM Confidence Estimation via Blind-Image Contrastive Ranking
by: Khanmohammadi, Reza, et al.
Published: (2026)
by: Khanmohammadi, Reza, et al.
Published: (2026)
How Reliable are Confidence Estimators for Large Reasoning Models? A Systematic Benchmark on High-Stakes Domains
by: Khanmohammadi, Reza, et al.
Published: (2026)
by: Khanmohammadi, Reza, et al.
Published: (2026)
The Influence of Biomedical Research on Future Business Funding: Analyzing Scientific Impact and Content in Industrial Investments
by: Khanmohammadi, Reza, et al.
Published: (2024)
by: Khanmohammadi, Reza, et al.
Published: (2024)
Grounding LLM Reasoning with Knowledge Graphs
by: Amayuelas, Alfonso, et al.
Published: (2025)
by: Amayuelas, Alfonso, et al.
Published: (2025)
FinQAPT: Empowering Financial Decisions with End-to-End LLM-driven Question Answering Pipeline
by: Singh, Kuldeep, et al.
Published: (2024)
by: Singh, Kuldeep, et al.
Published: (2024)
Conservative Bias in Large Language Models: Measuring Relation Predictions
by: Aguda, Toyin, et al.
Published: (2025)
by: Aguda, Toyin, et al.
Published: (2025)
FinNLI: Novel Dataset for Multi-Genre Financial Natural Language Inference Benchmarking
by: Magomere, Jabez, et al.
Published: (2025)
by: Magomere, Jabez, et al.
Published: (2025)
A Variational Approach for Mitigating Entity Bias in Relation Extraction
by: Mensah, Samuel, et al.
Published: (2025)
by: Mensah, Samuel, et al.
Published: (2025)
Large Language Models as Financial Data Annotators: A Study on Effectiveness and Efficiency
by: Aguda, Toyin, et al.
Published: (2024)
by: Aguda, Toyin, et al.
Published: (2024)
Iterative Prompt Refinement for Radiation Oncology Symptom Extraction Using Teacher-Student Large Language Models
by: Khanmohammadi, Reza, et al.
Published: (2024)
by: Khanmohammadi, Reza, et al.
Published: (2024)
Distill and Align Decomposition for Enhanced Claim Verification
by: Magomere, Jabez, et al.
Published: (2026)
by: Magomere, Jabez, et al.
Published: (2026)
Hybrid Student-Teacher Large Language Model Refinement for Cancer Toxicity Symptom Extraction
by: Khanmohammadi, Reza, et al.
Published: (2024)
by: Khanmohammadi, Reza, et al.
Published: (2024)
AI Analyst: Framework and Comprehensive Evaluation of Large Language Models for Financial Time Series Report Generation
by: Fons, Elizabeth, et al.
Published: (2025)
by: Fons, Elizabeth, et al.
Published: (2025)
Understanding and Exploiting Weight Update Sparsity for Communication-Efficient Distributed RL
by: Miahi, Erfan, et al.
Published: (2026)
by: Miahi, Erfan, et al.
Published: (2026)
Deep FinResearch Bench: Evaluating AI's Ability to Conduct Professional Financial Investment Research
by: Haque, Mirazul, et al.
Published: (2026)
by: Haque, Mirazul, et al.
Published: (2026)
WhisperPipe: A Resource-Efficient Streaming Architecture for Real-Time Automatic Speech Recognition
by: Ramezani, Erfan, et al.
Published: (2026)
by: Ramezani, Erfan, et al.
Published: (2026)
Evaluating and Calibrating LLM Confidence on Questions with Multiple Correct Answers
by: Wang, Yuhan, et al.
Published: (2026)
by: Wang, Yuhan, et al.
Published: (2026)
Process Supervision of Confidence Margin for Calibrated LLM Reasoning
by: Wang, Liaoyaqi, et al.
Published: (2026)
by: Wang, Liaoyaqi, et al.
Published: (2026)
WorldMemArena: Evaluating Multimodal Agent Memory Through Action-World Interaction
by: Liu, Chengzhi, et al.
Published: (2026)
by: Liu, Chengzhi, et al.
Published: (2026)
Interpretable LLM-based Table Question Answering
by: Nguyen, Giang, et al.
Published: (2024)
by: Nguyen, Giang, et al.
Published: (2024)
DocLLM: A layout-aware generative language model for multimodal document understanding
by: Wang, Dongsheng, et al.
Published: (2023)
by: Wang, Dongsheng, et al.
Published: (2023)
CritiCal: Can Critique Help LLM Uncertainty or Confidence Calibration?
by: Zong, Qing, et al.
Published: (2025)
by: Zong, Qing, et al.
Published: (2025)
Self-Aware Knowledge Probing: Evaluating Language Models' Relational Knowledge through Confidence Calibration
by: Kissling, Christopher, et al.
Published: (2026)
by: Kissling, Christopher, et al.
Published: (2026)
Agentic Confidence Calibration
by: Zhang, Jiaxin, et al.
Published: (2026)
by: Zhang, Jiaxin, et al.
Published: (2026)
Confidence over Time: Confidence Calibration with Temporal Logic for Large Language Model Reasoning
by: Mao, Zhenjiang, et al.
Published: (2026)
by: Mao, Zhenjiang, et al.
Published: (2026)
Automated stereotactic radiosurgery planning using a human-in-the-loop reasoning large language model agent
by: Nusrat, Humza, et al.
Published: (2025)
by: Nusrat, Humza, et al.
Published: (2025)
Autonomous Radiotherapy Treatment Planning Using DOLA: A Privacy-Preserving, LLM-Based Optimization Agent
by: Nusrat, Humza, et al.
Published: (2025)
by: Nusrat, Humza, et al.
Published: (2025)
Intertextual Parallel Detection in Biblical Hebrew: A Transformer-Based Benchmark
by: Smiley, David M.
Published: (2025)
by: Smiley, David M.
Published: (2025)
What Language Models Know But Don't Say: Non-Generative Prior Extraction for Generalization
by: Rezaeimanesh, Sara, et al.
Published: (2026)
by: Rezaeimanesh, Sara, et al.
Published: (2026)
Robustness of Transformer-Based Fluence Map Prediction Under Clinically Realistic Perturbations
by: Mgboh, Ujunwa, et al.
Published: (2026)
by: Mgboh, Ujunwa, et al.
Published: (2026)
When Can We Trust LLM Graders? Calibrating Confidence for Automated Assessment
by: Ferrer, Robinson, et al.
Published: (2026)
by: Ferrer, Robinson, et al.
Published: (2026)
OptimAI: Optimization from Natural Language Using LLM-Powered AI Agents
by: Thind, Raghav, et al.
Published: (2025)
by: Thind, Raghav, et al.
Published: (2025)
Calibrated Confidence Expression for Radiology Report Generation
by: Bani-Harouni, David, et al.
Published: (2026)
by: Bani-Harouni, David, et al.
Published: (2026)
Judging with Confidence: Calibrating Autoraters to Preference Distributions
by: Li, Zhuohang, et al.
Published: (2025)
by: Li, Zhuohang, et al.
Published: (2025)
Calibrating the Confidence of Large Language Models by Eliciting Fidelity
by: Zhang, Mozhi, et al.
Published: (2024)
by: Zhang, Mozhi, et al.
Published: (2024)
RepMatch: Quantifying Cross-Instance Similarities in Representation Space
by: Modarres, Mohammad Reza, et al.
Published: (2024)
by: Modarres, Mohammad Reza, et al.
Published: (2024)
ThoughtProbe: Classifier-Guided LLM Thought Space Exploration via Probing Representations
by: Wang, Zijian, et al.
Published: (2025)
by: Wang, Zijian, et al.
Published: (2025)
QA-Calibration of Language Model Confidence Scores
by: Manggala, Putra, et al.
Published: (2024)
by: Manggala, Putra, et al.
Published: (2024)
Calibrating Verbalized Confidence with Self-Generated Distractors
by: Wang, Victor, et al.
Published: (2025)
by: Wang, Victor, et al.
Published: (2025)
Fact-Level Confidence Calibration and Self-Correction
by: Yuan, Yige, et al.
Published: (2024)
by: Yuan, Yige, et al.
Published: (2024)
Similar Items
-
Grounded or Guessing? LVLM Confidence Estimation via Blind-Image Contrastive Ranking
by: Khanmohammadi, Reza, et al.
Published: (2026) -
How Reliable are Confidence Estimators for Large Reasoning Models? A Systematic Benchmark on High-Stakes Domains
by: Khanmohammadi, Reza, et al.
Published: (2026) -
The Influence of Biomedical Research on Future Business Funding: Analyzing Scientific Impact and Content in Industrial Investments
by: Khanmohammadi, Reza, et al.
Published: (2024) -
Grounding LLM Reasoning with Knowledge Graphs
by: Amayuelas, Alfonso, et al.
Published: (2025) -
FinQAPT: Empowering Financial Decisions with End-to-End LLM-driven Question Answering Pipeline
by: Singh, Kuldeep, et al.
Published: (2024)