Interpreting and Mitigating Unwanted Uncertainty in LLMs
Fuente:
arXiv
Salvato in:
| Autori principali: | Roy, Tiasa Singha, Jhaveri, Ayush Rajesh, Triantafyllopoulos, Ilias |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Can LLMs $\textit{understand}$ Math? -- Exploring the Pitfalls in Mathematical Reasoning
di: Roy, Tiasa Singha, et al.
Pubblicazione: (2025)
di: Roy, Tiasa Singha, et al.
Pubblicazione: (2025)
Failing to Falsify: Evaluating and Mitigating Confirmation Bias in Language Models
di: Jhaveri, Ayush Rajesh, et al.
Pubblicazione: (2026)
di: Jhaveri, Ayush Rajesh, et al.
Pubblicazione: (2026)
Neural Neural Scaling Laws
di: Hu, Michael Y., et al.
Pubblicazione: (2026)
di: Hu, Michael Y., et al.
Pubblicazione: (2026)
CREATE: Testing LLMs for Associative Creativity
di: Wadhwa, Manya, et al.
Pubblicazione: (2026)
di: Wadhwa, Manya, et al.
Pubblicazione: (2026)
Learning to Pay Attention: Unsupervised Modeling of Attentive and Inattentive Respondents in Survey Data
di: Triantafyllopoulos, Ilias, et al.
Pubblicazione: (2026)
di: Triantafyllopoulos, Ilias, et al.
Pubblicazione: (2026)
Conceptors for Semantic Steering
di: Triantafyllopoulos, Ilias, et al.
Pubblicazione: (2026)
di: Triantafyllopoulos, Ilias, et al.
Pubblicazione: (2026)
The Confidence Trap: Gender Bias and Predictive Certainty in LLMs
di: Sabir, Ahmed, et al.
Pubblicazione: (2026)
di: Sabir, Ahmed, et al.
Pubblicazione: (2026)
Detecting AI Hallucinations in Finance: An Information-Theoretic Method Cuts Hallucination Rate by 92%
di: Singha, Mainak
Pubblicazione: (2025)
di: Singha, Mainak
Pubblicazione: (2025)
BTZSC: A Benchmark for Zero-Shot Text Classification Across Cross-Encoders, Embedding Models, Rerankers and LLMs
di: Aarab, Ilias
Pubblicazione: (2026)
di: Aarab, Ilias
Pubblicazione: (2026)
MedLM: Exploring Language Models for Medical Question Answering Systems
di: Yagnik, Niraj, et al.
Pubblicazione: (2024)
di: Yagnik, Niraj, et al.
Pubblicazione: (2024)
Data Science with LLMs and Interpretable Models
di: Bordt, Sebastian, et al.
Pubblicazione: (2024)
di: Bordt, Sebastian, et al.
Pubblicazione: (2024)
A Concise Review of Hallucinations in LLMs and their Mitigation
di: Pulkundwar, Parth, et al.
Pubblicazione: (2025)
di: Pulkundwar, Parth, et al.
Pubblicazione: (2025)
Uncertainty as Feature Gaps: Epistemic Uncertainty Quantification of LLMs in Contextual Question-Answering
di: Bakman, Yavuz, et al.
Pubblicazione: (2025)
di: Bakman, Yavuz, et al.
Pubblicazione: (2025)
UCD: Unlearning in LLMs via Contrastive Decoding
di: Suriyakumar, Vinith M., et al.
Pubblicazione: (2025)
di: Suriyakumar, Vinith M., et al.
Pubblicazione: (2025)
GRU: Mitigating the Trade-off between Unlearning and Retention for LLMs
di: Wang, Yue, et al.
Pubblicazione: (2025)
di: Wang, Yue, et al.
Pubblicazione: (2025)
On Pruning State-Space LLMs
di: Ghattas, Tamer, et al.
Pubblicazione: (2025)
di: Ghattas, Tamer, et al.
Pubblicazione: (2025)
The Illusion of Certainty: Uncertainty Quantification for LLMs Fails under Ambiguity
di: Tomov, Tim, et al.
Pubblicazione: (2025)
di: Tomov, Tim, et al.
Pubblicazione: (2025)
MetaFaith: Faithful Natural Language Uncertainty Expression in LLMs
di: Liu, Gabrielle Kaili-May, et al.
Pubblicazione: (2025)
di: Liu, Gabrielle Kaili-May, et al.
Pubblicazione: (2025)
Beyond the Score: Uncertainty-Calibrated LLMs for Automated Essay Assessment
di: Karim, Ahmed, et al.
Pubblicazione: (2025)
di: Karim, Ahmed, et al.
Pubblicazione: (2025)
ALIEN: Aligned Entropy Head for Improving Uncertainty Estimation of LLMs
di: Zabolotnyi, Artem, et al.
Pubblicazione: (2025)
di: Zabolotnyi, Artem, et al.
Pubblicazione: (2025)
Uncovering Gaps in How Humans and LLMs Interpret Subjective Language
di: Jones, Erik, et al.
Pubblicazione: (2025)
di: Jones, Erik, et al.
Pubblicazione: (2025)
Understanding and Mitigating the Uncertainty in Zero-Shot Translation
di: Wang, Wenxuan, et al.
Pubblicazione: (2022)
di: Wang, Wenxuan, et al.
Pubblicazione: (2022)
ThinkEdit: Interpretable Weight Editing to Mitigate Overly Short Thinking in Reasoning Models
di: Sun, Chung-En, et al.
Pubblicazione: (2025)
di: Sun, Chung-En, et al.
Pubblicazione: (2025)
ZClip: Adaptive Spike Mitigation for LLM Pre-Training
di: Kumar, Abhay, et al.
Pubblicazione: (2025)
di: Kumar, Abhay, et al.
Pubblicazione: (2025)
Mitigating Quantization Errors Due to Activation Spikes in GLU-Based LLMs
di: Yang, Jaewoo, et al.
Pubblicazione: (2024)
di: Yang, Jaewoo, et al.
Pubblicazione: (2024)
Can Performant LLMs Be Ethical? Quantifying the Impact of Web Crawling Opt-Outs
di: Fan, Dongyang, et al.
Pubblicazione: (2025)
di: Fan, Dongyang, et al.
Pubblicazione: (2025)
Interpreting the Effects of Quantization on LLMs
di: Singh, Manpreet, et al.
Pubblicazione: (2025)
di: Singh, Manpreet, et al.
Pubblicazione: (2025)
MARS: Meaning-Aware Response Scoring for Uncertainty Estimation in Generative LLMs
di: Bakman, Yavuz Faruk, et al.
Pubblicazione: (2024)
di: Bakman, Yavuz Faruk, et al.
Pubblicazione: (2024)
LLMs Explain't: A Post-Mortem on Semantic Interpretability in Transformer Models
di: Abdelhalim, Alhassan, et al.
Pubblicazione: (2026)
di: Abdelhalim, Alhassan, et al.
Pubblicazione: (2026)
LANCET: Neural Intervention via Structural Entropy for Mitigating Faithfulness Hallucinations in LLMs
di: Wang, Chenxu, et al.
Pubblicazione: (2026)
di: Wang, Chenxu, et al.
Pubblicazione: (2026)
CoT-UQ: Improving Response-wise Uncertainty Quantification in LLMs with Chain-of-Thought
di: Zhang, Boxuan, et al.
Pubblicazione: (2025)
di: Zhang, Boxuan, et al.
Pubblicazione: (2025)
LLMs Should Express Uncertainty Explicitly
di: Guo, Junyu, et al.
Pubblicazione: (2026)
di: Guo, Junyu, et al.
Pubblicazione: (2026)
OncoReason: Structuring Clinical Reasoning in LLMs for Robust and Interpretable Survival Prediction
di: Hemadri, Raghu Vamshi, et al.
Pubblicazione: (2025)
di: Hemadri, Raghu Vamshi, et al.
Pubblicazione: (2025)
Mitigating Semantic Drift: Evaluating LLMs' Efficacy in Psychotherapy through MI Dialogue Summarization
di: Kumar, Vivek, et al.
Pubblicazione: (2025)
di: Kumar, Vivek, et al.
Pubblicazione: (2025)
Risks, Causes, and Mitigations of Widespread Deployments of Large Language Models (LLMs): A Survey
di: Sakib, Md Nazmus, et al.
Pubblicazione: (2024)
di: Sakib, Md Nazmus, et al.
Pubblicazione: (2024)
The Landscape of Memorization in LLMs: Mechanisms, Measurement, and Mitigation
di: Xiong, Alexander, et al.
Pubblicazione: (2025)
di: Xiong, Alexander, et al.
Pubblicazione: (2025)
LLMs as Implicit Imputers: Uncertainty Should Scale with Missing Information
di: van Buuren, Stef
Pubblicazione: (2026)
di: van Buuren, Stef
Pubblicazione: (2026)
DICTDIS: Dictionary Constrained Disambiguation for Improved NMT
di: Maheshwari, Ayush, et al.
Pubblicazione: (2022)
di: Maheshwari, Ayush, et al.
Pubblicazione: (2022)
StreetMath: Study of LLMs' Approximation Behaviors
di: Tseng, Chiung-Yi, et al.
Pubblicazione: (2025)
di: Tseng, Chiung-Yi, et al.
Pubblicazione: (2025)
TeleLoRA: Teleporting Model-Specific Alignment Across LLMs
di: Lin, Xiao, et al.
Pubblicazione: (2025)
di: Lin, Xiao, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Can LLMs $\textit{understand}$ Math? -- Exploring the Pitfalls in Mathematical Reasoning
di: Roy, Tiasa Singha, et al.
Pubblicazione: (2025) -
Failing to Falsify: Evaluating and Mitigating Confirmation Bias in Language Models
di: Jhaveri, Ayush Rajesh, et al.
Pubblicazione: (2026) -
Neural Neural Scaling Laws
di: Hu, Michael Y., et al.
Pubblicazione: (2026) -
CREATE: Testing LLMs for Associative Creativity
di: Wadhwa, Manya, et al.
Pubblicazione: (2026) -
Learning to Pay Attention: Unsupervised Modeling of Attentive and Inattentive Respondents in Survey Data
di: Triantafyllopoulos, Ilias, et al.
Pubblicazione: (2026)