On the Robustness of Verbal Confidence of LLMs in Adversarial Attacks
Fuente:
arXiv
Saved in:
| Main Authors: | Obadinma, Stephen, Zhu, Xiaodan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Calibration Attacks: A Comprehensive Study of Adversarial Attacks on Model Confidence
by: Obadinma, Stephen, et al.
Published: (2024)
by: Obadinma, Stephen, et al.
Published: (2024)
RedDebate: Safer Responses Through Multi-Agent Red Teaming Debates
by: Asad, Ali, et al.
Published: (2025)
by: Asad, Ali, et al.
Published: (2025)
On Verbalized Confidence Scores for LLMs
by: Yang, Daniel, et al.
Published: (2024)
by: Yang, Daniel, et al.
Published: (2024)
NAACL: Noise-AwAre Verbal Confidence Calibration for Robust LLMs in RAG Systems
by: Liu, Jiayu, et al.
Published: (2026)
by: Liu, Jiayu, et al.
Published: (2026)
Influential Training Data Retrieval for Explaining Verbalized Confidence of LLMs
by: Xia, Yuxi, et al.
Published: (2026)
by: Xia, Yuxi, et al.
Published: (2026)
How do LLMs Compute Verbal Confidence
by: Kumaran, Dharshan, et al.
Published: (2026)
by: Kumaran, Dharshan, et al.
Published: (2026)
Direct Confidence Alignment: Aligning Verbalized Confidence with Internal Confidence In Large Language Models
by: Zhang, Glenn, et al.
Published: (2025)
by: Zhang, Glenn, et al.
Published: (2025)
Wired for Overconfidence: A Mechanistic Perspective on Inflated Verbalized Confidence in LLMs
by: Zhao, Tianyi, et al.
Published: (2026)
by: Zhao, Tianyi, et al.
Published: (2026)
ADVICE: Answer-Dependent Verbalized Confidence Estimation
by: Seo, Ki Jung, et al.
Published: (2025)
by: Seo, Ki Jung, et al.
Published: (2025)
Are LLM Decisions Faithful to Verbal Confidence?
by: Wang, Jiawei, et al.
Published: (2026)
by: Wang, Jiawei, et al.
Published: (2026)
NatLogAttack: A Framework for Attacking Natural Language Inference Models with Natural Logic
by: Zheng, Zi'ou, et al.
Published: (2023)
by: Zheng, Zi'ou, et al.
Published: (2023)
Calibrating Verbalized Confidence with Self-Generated Distractors
by: Wang, Victor, et al.
Published: (2025)
by: Wang, Victor, et al.
Published: (2025)
Are Large Language Models More Honest in Their Probabilistic or Verbalized Confidence?
by: Ni, Shiyu, et al.
Published: (2024)
by: Ni, Shiyu, et al.
Published: (2024)
LoVeC: Reinforcement Learning for Better Verbalized Confidence in Long-Form Generations
by: Zhang, Caiqi, et al.
Published: (2025)
by: Zhang, Caiqi, et al.
Published: (2025)
Adversarial Attack for Explanation Robustness of Rationalization Models
by: Zhang, Yuankai, et al.
Published: (2024)
by: Zhang, Yuankai, et al.
Published: (2024)
Let the Model Distribute Its Doubt: Confidence Estimation through Verbalized Probability Distribution
by: Wang, Ante, et al.
Published: (2025)
by: Wang, Ante, et al.
Published: (2025)
ConfTuner: Training Large Language Models to Express Their Confidence Verbally
by: Li, Yibo, et al.
Published: (2025)
by: Li, Yibo, et al.
Published: (2025)
ORCE: Order-Aware Alignment of Verbalized Confidence in Large Language Models
by: Li, Chen, et al.
Published: (2026)
by: Li, Chen, et al.
Published: (2026)
Robustness of Large Language Models Against Adversarial Attacks
by: Tao, Yiyi, et al.
Published: (2024)
by: Tao, Yiyi, et al.
Published: (2024)
Factual Confidence of LLMs: on Reliability and Robustness of Current Estimators
by: Mahaut, Matéo, et al.
Published: (2024)
by: Mahaut, Matéo, et al.
Published: (2024)
Robust Utility-Preserving Text Anonymization Based on Large Language Models
by: Yang, Tianyu, et al.
Published: (2024)
by: Yang, Tianyu, et al.
Published: (2024)
Verbal Confidence Saturation in 3-9B Open-Weight Instruction-Tuned LLMs: A Pre-Registered Psychometric Validity Screen
by: Cacioli, Jon-Paul
Published: (2026)
by: Cacioli, Jon-Paul
Published: (2026)
Confidence Estimation for LLMs in Multi-turn Interactions
by: Zhang, Caiqi, et al.
Published: (2026)
by: Zhang, Caiqi, et al.
Published: (2026)
Verbalizing LLMs' assumptions to explain and control sycophancy
by: Cheng, Myra, et al.
Published: (2026)
by: Cheng, Myra, et al.
Published: (2026)
Verbalized Confidence Triggers Self-Verification: Emergent Behavior Without Explicit Reasoning Supervision
by: Jang, Chaeyun, et al.
Published: (2025)
by: Jang, Chaeyun, et al.
Published: (2025)
Robustness of Misinformation Classification Systems to Adversarial Examples Through BeamAttack
by: Fazla, Arnisa, et al.
Published: (2025)
by: Fazla, Arnisa, et al.
Published: (2025)
DiffuseDef: Improved Robustness to Adversarial Attacks via Iterative Denoising
by: Li, Zhenhao, et al.
Published: (2024)
by: Li, Zhenhao, et al.
Published: (2024)
InternalInspector $I^2$: Robust Confidence Estimation in LLMs through Internal States
by: Beigi, Mohammad, et al.
Published: (2024)
by: Beigi, Mohammad, et al.
Published: (2024)
LimeAttack: Local Explainable Method for Textual Hard-Label Adversarial Attack
by: Zhu, Hai, et al.
Published: (2023)
by: Zhu, Hai, et al.
Published: (2023)
CARES: Comprehensive Evaluation of Safety and Adversarial Robustness in Medical LLMs
by: Chen, Sijia, et al.
Published: (2025)
by: Chen, Sijia, et al.
Published: (2025)
Reasoning Robustness of LLMs to Adversarial Typographical Errors
by: Gan, Esther, et al.
Published: (2024)
by: Gan, Esther, et al.
Published: (2024)
Causal Front-Door Adjustment for Robust Jailbreak Attacks on LLMs
by: Zhou, Yao, et al.
Published: (2026)
by: Zhou, Yao, et al.
Published: (2026)
Adversarial Tuning: Defending Against Jailbreak Attacks for LLMs
by: Liu, Fan, et al.
Published: (2024)
by: Liu, Fan, et al.
Published: (2024)
Fine-grained Verbal Attack Detection via a Hierarchical Divide-and-Conquer Framework
by: Zheng, Quan, et al.
Published: (2026)
by: Zheng, Quan, et al.
Published: (2026)
Mitigating Adversarial Attacks in LLMs through Defensive Suffix Generation
by: Kim, Minkyoung, et al.
Published: (2024)
by: Kim, Minkyoung, et al.
Published: (2024)
Code Prompting Elicits Conditional Reasoning Abilities in Text+Code LLMs
by: Puerto, Haritz, et al.
Published: (2024)
by: Puerto, Haritz, et al.
Published: (2024)
Multicalibration for Confidence Scoring in LLMs
by: Detommaso, Gianluca, et al.
Published: (2024)
by: Detommaso, Gianluca, et al.
Published: (2024)
Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs
by: Sheshadri, Abhay, et al.
Published: (2024)
by: Sheshadri, Abhay, et al.
Published: (2024)
Universal Acoustic Adversarial Attacks for Flexible Control of Speech-LLMs
by: Ma, Rao, et al.
Published: (2025)
by: Ma, Rao, et al.
Published: (2025)
Self-Evaluation as a Defense Against Adversarial Attacks on LLMs
by: Brown, Hannah, et al.
Published: (2024)
by: Brown, Hannah, et al.
Published: (2024)
Similar Items
-
Calibration Attacks: A Comprehensive Study of Adversarial Attacks on Model Confidence
by: Obadinma, Stephen, et al.
Published: (2024) -
RedDebate: Safer Responses Through Multi-Agent Red Teaming Debates
by: Asad, Ali, et al.
Published: (2025) -
On Verbalized Confidence Scores for LLMs
by: Yang, Daniel, et al.
Published: (2024) -
NAACL: Noise-AwAre Verbal Confidence Calibration for Robust LLMs in RAG Systems
by: Liu, Jiayu, et al.
Published: (2026) -
Influential Training Data Retrieval for Explaining Verbalized Confidence of LLMs
by: Xia, Yuxi, et al.
Published: (2026)