Does Alignment Tuning Really Break LLMs' Internal Confidence?
Fuente:
arXiv
Saved in:
| Main Authors: | Oh, Hongseok, Hwang, Wonseok |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Do Vision Encoders Truly Explain Object Hallucination?: Mitigating Object Hallucination via Simple Fine-Grained CLIPScore
by: Oh, Hongseok, et al.
Published: (2025)
by: Oh, Hongseok, et al.
Published: (2025)
Korean Canonical Legal Benchmark: Toward Knowledge-Independent Evaluation of LLMs' Legal Reasoning Capabilities
by: Oh, Hongseok, et al.
Published: (2025)
by: Oh, Hongseok, et al.
Published: (2025)
Direct Alignment of Draft Model for Speculative Decoding with Chat-Fine-Tuned LLMs
by: Goel, Raghavv, et al.
Published: (2024)
by: Goel, Raghavv, et al.
Published: (2024)
Do LLMs Really Forget? Evaluating Unlearning with Knowledge Correlation and Confidence Awareness
by: Wei, Rongzhe, et al.
Published: (2025)
by: Wei, Rongzhe, et al.
Published: (2025)
LRAGE: Legal Retrieval Augmented Generation Evaluation Tool
by: Park, Minhu, et al.
Published: (2025)
by: Park, Minhu, et al.
Published: (2025)
Show Your Work with Confidence: Confidence Bands for Tuning Curves
by: Lourie, Nicholas, et al.
Published: (2023)
by: Lourie, Nicholas, et al.
Published: (2023)
Enhancing Automated Essay Scoring with Three Techniques: Two-Stage Fine-Tuning, Score Alignment, and Self-Training
by: Choi, Hongseok, et al.
Published: (2026)
by: Choi, Hongseok, et al.
Published: (2026)
Multicalibration for Confidence Scoring in LLMs
by: Detommaso, Gianluca, et al.
Published: (2024)
by: Detommaso, Gianluca, et al.
Published: (2024)
SteerConf: Steering LLMs for Confidence Elicitation
by: Zhou, Ziang, et al.
Published: (2025)
by: Zhou, Ziang, et al.
Published: (2025)
Single Character Perturbations Break LLM Alignment
by: Lin, Leon, et al.
Published: (2024)
by: Lin, Leon, et al.
Published: (2024)
On the Consideration of AI Openness: Can Good Intent Be Abused?
by: Kim, Yeeun, et al.
Published: (2024)
by: Kim, Yeeun, et al.
Published: (2024)
SelfReflect: Can LLMs Communicate Their Internal Answer Distribution?
by: Kirchhof, Michael, et al.
Published: (2025)
by: Kirchhof, Michael, et al.
Published: (2025)
AdaEDL: Early Draft Stopping for Speculative Decoding of Large Language Models via an Entropy-based Lower Bound on Token Acceptance Probability
by: Agrawal, Sudhanshu, et al.
Published: (2024)
by: Agrawal, Sudhanshu, et al.
Published: (2024)
Factual Confidence of LLMs: on Reliability and Robustness of Current Estimators
by: Mahaut, Matéo, et al.
Published: (2024)
by: Mahaut, Matéo, et al.
Published: (2024)
The Confidence Trap: Gender Bias and Predictive Certainty in LLMs
by: Sabir, Ahmed, et al.
Published: (2026)
by: Sabir, Ahmed, et al.
Published: (2026)
When Style Breaks Safety: Defending LLMs Against Superficial Style Alignment
by: Xiao, Yuxin, et al.
Published: (2025)
by: Xiao, Yuxin, et al.
Published: (2025)
Does LLM Alignment Really Need Diversity? An Empirical Study of Adapting RLVR Methods for Moral Reasoning
by: Zhang, Zhaowei, et al.
Published: (2026)
by: Zhang, Zhaowei, et al.
Published: (2026)
ORCE: Order-Aware Alignment of Verbalized Confidence in Large Language Models
by: Li, Chen, et al.
Published: (2026)
by: Li, Chen, et al.
Published: (2026)
Confidence Under the Hood: An Investigation into the Confidence-Probability Alignment in Large Language Models
by: Kumar, Abhishek, et al.
Published: (2024)
by: Kumar, Abhishek, et al.
Published: (2024)
Confidence-Credibility Aware Weighted Ensembles of Small LLMs Outperform Large LLMs in Emotion Detection
by: Elgabry, Menna, et al.
Published: (2025)
by: Elgabry, Menna, et al.
Published: (2025)
MICE for CATs: Model-Internal Confidence Estimation for Calibrating Agents with Tools
by: Subramani, Nishant, et al.
Published: (2025)
by: Subramani, Nishant, et al.
Published: (2025)
Learning to Route LLMs with Confidence Tokens
by: Chuang, Yu-Neng, et al.
Published: (2024)
by: Chuang, Yu-Neng, et al.
Published: (2024)
PCS: Perceived Confidence Scoring of Black Box LLMs with Metamorphic Relations
by: Salimian, Sina, et al.
Published: (2025)
by: Salimian, Sina, et al.
Published: (2025)
ConfClip: Confidence-Weighted and Clipped Reward for Reinforcement Learning in LLMs
by: Zhang, Bonan, et al.
Published: (2025)
by: Zhang, Bonan, et al.
Published: (2025)
Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models
by: Li, Pengyi, et al.
Published: (2025)
by: Li, Pengyi, et al.
Published: (2025)
Calibration Collapse Under Sycophancy Fine-Tuning: How Reward Hacking Breaks Uncertainty Quantification in LLMs
by: Sahoo, Subramanyam
Published: (2026)
by: Sahoo, Subramanyam
Published: (2026)
How Does Quantization Affect Multilingual LLMs?
by: Marchisio, Kelly, et al.
Published: (2024)
by: Marchisio, Kelly, et al.
Published: (2024)
Re-Emergent Misalignment: How Narrow Fine-Tuning Erodes Safety Alignment in LLMs
by: Giordani, Jeremiah
Published: (2025)
by: Giordani, Jeremiah
Published: (2025)
Picky LLMs and Unreliable RMs: An Empirical Study on Safety Alignment after Instruction Tuning
by: Li, Guanlin, et al.
Published: (2025)
by: Li, Guanlin, et al.
Published: (2025)
NeuronTune: Fine-Grained Neuron Modulation for Balanced Safety-Utility Alignment in LLMs
by: Pan, Birong, et al.
Published: (2025)
by: Pan, Birong, et al.
Published: (2025)
Reasoning Boosts Opinion Alignment in LLMs
by: Berdoz, Frédéric, et al.
Published: (2026)
by: Berdoz, Frédéric, et al.
Published: (2026)
STBLLM: Breaking the 1-Bit Barrier with Structured Binary LLMs
by: Dong, Peijie, et al.
Published: (2024)
by: Dong, Peijie, et al.
Published: (2024)
How do LLMs Compute Verbal Confidence
by: Kumaran, Dharshan, et al.
Published: (2026)
by: Kumaran, Dharshan, et al.
Published: (2026)
House of Cards: Massive Weights in LLMs
by: Oh, Jaehoon, et al.
Published: (2024)
by: Oh, Jaehoon, et al.
Published: (2024)
QEFT: Quantization for Efficient Fine-Tuning of LLMs
by: Lee, Changhun, et al.
Published: (2024)
by: Lee, Changhun, et al.
Published: (2024)
Internal Chain-of-Thought: Empirical Evidence for Layer-wise Subtask Scheduling in LLMs
by: Yang, Zhipeng, et al.
Published: (2025)
by: Yang, Zhipeng, et al.
Published: (2025)
Entropy-Adaptive Fine-Tuning: Resolving Confident Conflicts to Mitigate Forgetting
by: Diao, Muxi, et al.
Published: (2026)
by: Diao, Muxi, et al.
Published: (2026)
Iterative Self-Tuning LLMs for Enhanced Jailbreaking Capabilities
by: Sun, Chung-En, et al.
Published: (2024)
by: Sun, Chung-En, et al.
Published: (2024)
Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?
by: Kim, Jeonghye, et al.
Published: (2026)
by: Kim, Jeonghye, et al.
Published: (2026)
Towards Robust and Parameter-Efficient Knowledge Unlearning for LLMs
by: Cha, Sungmin, et al.
Published: (2024)
by: Cha, Sungmin, et al.
Published: (2024)
Similar Items
-
Do Vision Encoders Truly Explain Object Hallucination?: Mitigating Object Hallucination via Simple Fine-Grained CLIPScore
by: Oh, Hongseok, et al.
Published: (2025) -
Korean Canonical Legal Benchmark: Toward Knowledge-Independent Evaluation of LLMs' Legal Reasoning Capabilities
by: Oh, Hongseok, et al.
Published: (2025) -
Direct Alignment of Draft Model for Speculative Decoding with Chat-Fine-Tuned LLMs
by: Goel, Raghavv, et al.
Published: (2024) -
Do LLMs Really Forget? Evaluating Unlearning with Knowledge Correlation and Confidence Awareness
by: Wei, Rongzhe, et al.
Published: (2025) -
LRAGE: Legal Retrieval Augmented Generation Evaluation Tool
by: Park, Minhu, et al.
Published: (2025)