InternalInspector $I^2$: Robust Confidence Estimation in LLMs through Internal States
Fuente:
arXiv
Saved in:
| Main Authors: | Beigi, Mohammad, Shen, Ying, Yang, Runing, Lin, Zihao, Wang, Qifan, Mohan, Ankith, He, Jianfeng, Jin, Ming, Lu, Chang-Tien, Huang, Lifu |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Adversarial Reward Auditing for Active Detection and Mitigation of Reward Hacking
by: Beigi, Mohammad, et al.
Published: (2026)
by: Beigi, Mohammad, et al.
Published: (2026)
IR$^3$: Contrastive Inverse Reinforcement Learning for Interpretable Detection and Mitigation of Reward Hacking
by: Beigi, Mohammad, et al.
Published: (2026)
by: Beigi, Mohammad, et al.
Published: (2026)
Sycophancy Mitigation Through Reinforcement Learning with Uncertainty-Aware Adaptive Reasoning Trajectories
by: Beigi, Mohammad, et al.
Published: (2025)
by: Beigi, Mohammad, et al.
Published: (2025)
Rethinking the Uncertainty: A Critical Review and Analysis in the Era of Large Language Models
by: Beigi, Mohammad, et al.
Published: (2024)
by: Beigi, Mohammad, et al.
Published: (2024)
Can We Trust the Performance Evaluation of Uncertainty Estimation Methods in Text Summarization?
by: He, Jianfeng, et al.
Published: (2024)
by: He, Jianfeng, et al.
Published: (2024)
Navigating the Dual Facets: A Comprehensive Evaluation of Sequential Memory Editing in Large Language Models
by: Lin, Zihao, et al.
Published: (2024)
by: Lin, Zihao, et al.
Published: (2024)
Does Alignment Tuning Really Break LLMs' Internal Confidence?
by: Oh, Hongseok, et al.
Published: (2024)
by: Oh, Hongseok, et al.
Published: (2024)
MICE for CATs: Model-Internal Confidence Estimation for Calibrating Agents with Tools
by: Subramani, Nishant, et al.
Published: (2025)
by: Subramani, Nishant, et al.
Published: (2025)
How LLMs Detect and Correct Their Own Errors: The Role of Internal Confidence Signals
by: Kumaran, Dharshan, et al.
Published: (2026)
by: Kumaran, Dharshan, et al.
Published: (2026)
A hybrid quantum-classical algorithm for Bayes-optimal quantum state discrimination using the source code
by: Mohan, Ankith, et al.
Published: (2023)
by: Mohan, Ankith, et al.
Published: (2023)
Anatomy of the Modality Gap: Dissecting the Internal States of End-to-End Speech LLMs
by: Hsu, Ming-Hao, et al.
Published: (2026)
by: Hsu, Ming-Hao, et al.
Published: (2026)
Direct Confidence Alignment: Aligning Verbalized Confidence with Internal Confidence In Large Language Models
by: Zhang, Glenn, et al.
Published: (2025)
by: Zhang, Glenn, et al.
Published: (2025)
Guiding a Diffusion Transformer with the Internal Dynamics of Itself
by: Zhou, Xingyu, et al.
Published: (2025)
by: Zhou, Xingyu, et al.
Published: (2025)
Enhancing Uncertainty Estimation in LLMs with Expectation of Aggregated Internal Belief
by: Xiao, Zeguan, et al.
Published: (2025)
by: Xiao, Zeguan, et al.
Published: (2025)
Enhancing Character-Level Understanding in LLMs through Token Internal Structure Learning
by: Xu, Zhu, et al.
Published: (2024)
by: Xu, Zhu, et al.
Published: (2024)
The pretty bad measurement
by: McIrvin, Caleb, et al.
Published: (2024)
by: McIrvin, Caleb, et al.
Published: (2024)
INSIDE: LLMs' Internal States Retain the Power of Hallucination Detection
by: Chen, Chao, et al.
Published: (2024)
by: Chen, Chao, et al.
Published: (2024)
Internal State Estimation in Groups via Active Information Gathering
by: Ji, Xuebo, et al.
Published: (2025)
by: Ji, Xuebo, et al.
Published: (2025)
AIRepr: An Analyst-Inspector Framework for Evaluating Reproducibility of LLMs in Data Science
by: Zeng, Qiuhai, et al.
Published: (2025)
by: Zeng, Qiuhai, et al.
Published: (2025)
LLM Braces: Straightening Out LLM Predictions with Relevant Sub-Updates
by: Shen, Ying, et al.
Published: (2025)
by: Shen, Ying, et al.
Published: (2025)
Quantum heuristics for linear optimization over large separable operators
by: Mohan, Ankith, et al.
Published: (2025)
by: Mohan, Ankith, et al.
Published: (2025)
Inside Out: Uncovering How Comment Internalization Steers LLMs for Better or Worse
by: Imani, Aaron, et al.
Published: (2025)
by: Imani, Aaron, et al.
Published: (2025)
Risk Assessment Framework for Code LLMs via Leveraging Internal States
by: Huang, Yuheng, et al.
Published: (2025)
by: Huang, Yuheng, et al.
Published: (2025)
On LLMs' Internal Representation of Code Correctness
by: Ribeiro, Francisco, et al.
Published: (2025)
by: Ribeiro, Francisco, et al.
Published: (2025)
Hallucination Detection with the Internal Layers of LLMs
by: Preiß, Martin
Published: (2025)
by: Preiß, Martin
Published: (2025)
Factual Confidence of LLMs: on Reliability and Robustness of Current Estimators
by: Mahaut, Matéo, et al.
Published: (2024)
by: Mahaut, Matéo, et al.
Published: (2024)
Multimodal Instruction Tuning with Conditional Mixture of LoRA
by: Shen, Ying, et al.
Published: (2024)
by: Shen, Ying, et al.
Published: (2024)
Confidence and Calibration of Activation Oracles for Reliable Interpretation of Language Model Internals
by: Torrielli, Federico, et al.
Published: (2026)
by: Torrielli, Federico, et al.
Published: (2026)
Transparentize the Internal and External Knowledge Utilization in LLMs with Trustworthy Citation
by: Shen, Jiajun, et al.
Published: (2025)
by: Shen, Jiajun, et al.
Published: (2025)
Beyond Surface Statistics: Robust Conformal Prediction for LLMs via Internal Representations
by: Wang, Yanli, et al.
Published: (2026)
by: Wang, Yanli, et al.
Published: (2026)
Do LLMs Experience an Internal Polylogue? Investigating Reasoning through the Lens of Personas
by: Herrmann, Nils A., et al.
Published: (2026)
by: Herrmann, Nils A., et al.
Published: (2026)
Safety Instincts: LLMs Learn to Trust Their Internal Compass for Self-Defense
by: Shen, Guobin, et al.
Published: (2025)
by: Shen, Guobin, et al.
Published: (2025)
Beyond Benchmarks: Understanding Mixture-of-Experts Models through Internal Mechanisms
by: Ying, Jiahao, et al.
Published: (2025)
by: Ying, Jiahao, et al.
Published: (2025)
Probing the Lack of Stable Internal Beliefs in LLMs
by: Luo, Yifan, et al.
Published: (2026)
by: Luo, Yifan, et al.
Published: (2026)
Error-driven Data-efficient Large Multimodal Model Tuning
by: Yao, Barry Menglong, et al.
Published: (2024)
by: Yao, Barry Menglong, et al.
Published: (2024)
Internal absolute geometry I: desingularization
by: Bohlen, Karsten
Published: (2021)
by: Bohlen, Karsten
Published: (2021)
Unsupervised Domain Adaptation Using Compact Internal Representations
by: Rostami, Mohammad
Published: (2024)
by: Rostami, Mohammad
Published: (2024)
github.com/FusionInspector/FusionInspector/fusion_inspector
by: Brian Haas
Published: (2026)
by: Brian Haas
Published: (2026)
github.com/FusionInspector/FusionInspector/fusion_inspector
by: Brian Haas
Published: (2026)
by: Brian Haas
Published: (2026)
How to Fix Left Colonic Anastomoses in Place to Prevent Internal Herniation
by: Simon Bennet, et al.
Published: (2025)
by: Simon Bennet, et al.
Published: (2025)
Similar Items
-
Adversarial Reward Auditing for Active Detection and Mitigation of Reward Hacking
by: Beigi, Mohammad, et al.
Published: (2026) -
IR$^3$: Contrastive Inverse Reinforcement Learning for Interpretable Detection and Mitigation of Reward Hacking
by: Beigi, Mohammad, et al.
Published: (2026) -
Sycophancy Mitigation Through Reinforcement Learning with Uncertainty-Aware Adaptive Reasoning Trajectories
by: Beigi, Mohammad, et al.
Published: (2025) -
Rethinking the Uncertainty: A Critical Review and Analysis in the Era of Large Language Models
by: Beigi, Mohammad, et al.
Published: (2024) -
Can We Trust the Performance Evaluation of Uncertainty Estimation Methods in Text Summarization?
by: He, Jianfeng, et al.
Published: (2024)