Wisdom is Knowing What not to Say: Hallucination-Free LLMs Unlearning via Attention Shifting
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Tan, Chenchen, Qu, Youyang, Li, Xinghao, Zhang, Hui, Cui, Shujie, Chen, Cunjian, Gao, Longxiang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Less is More: Geometric Unlearning for LLMs with Minimal Data Disclosure
von: Tan, Chenchen, et al.
Veröffentlicht: (2026)
von: Tan, Chenchen, et al.
Veröffentlicht: (2026)
When Harmless Words Harm: A New Threat to LLM Safety via Conceptual Triggers
von: Zhang, Zhaoxin, et al.
Veröffentlicht: (2025)
von: Zhang, Zhaoxin, et al.
Veröffentlicht: (2025)
Large Language Models Know What To Say But Not When To Speak
von: Umair, Muhammad, et al.
Veröffentlicht: (2024)
von: Umair, Muhammad, et al.
Veröffentlicht: (2024)
You Know What I'm Saying: Jailbreak Attack via Implicit Reference
von: Wu, Tianyu, et al.
Veröffentlicht: (2024)
von: Wu, Tianyu, et al.
Veröffentlicht: (2024)
LLMs Know More About Numbers than They Can Say
von: Yuchi, Fengting, et al.
Veröffentlicht: (2026)
von: Yuchi, Fengting, et al.
Veröffentlicht: (2026)
Recent Advances in Federated Learning Driven Large Language Models: A Survey on Architecture, Performance, and Security
von: Qu, Youyang, et al.
Veröffentlicht: (2024)
von: Qu, Youyang, et al.
Veröffentlicht: (2024)
Honest AI: Fine-Tuning "Small" Language Models to Say "I Don't Know", and Reducing Hallucination in RAG
von: Chen, Xinxi, et al.
Veröffentlicht: (2024)
von: Chen, Xinxi, et al.
Veröffentlicht: (2024)
What Models Know, How Well They Know It: Knowledge-Weighted Fine-Tuning for Learning When to Say "I Don't Know"
von: Lee, Joosung, et al.
Veröffentlicht: (2026)
von: Lee, Joosung, et al.
Veröffentlicht: (2026)
Unsupervised Layer-Wise Dynamic Test Time Adaptation for LLMs
von: Xu, Longhuan, et al.
Veröffentlicht: (2026)
von: Xu, Longhuan, et al.
Veröffentlicht: (2026)
Only Say What You Know: Calibration-Aware Generation for Long-Form Factuality
von: Luo, Wen, et al.
Veröffentlicht: (2026)
von: Luo, Wen, et al.
Veröffentlicht: (2026)
HearSay Benchmark: Do Audio LLMs Leak What They Hear?
von: Wang, Jin, et al.
Veröffentlicht: (2026)
von: Wang, Jin, et al.
Veröffentlicht: (2026)
When Models Know More Than They Say: Probing Analogical Reasoning in LLMs
von: McGovern, Hope, et al.
Veröffentlicht: (2026)
von: McGovern, Hope, et al.
Veröffentlicht: (2026)
What Language Models Know But Don't Say: Non-Generative Prior Extraction for Generalization
von: Rezaeimanesh, Sara, et al.
Veröffentlicht: (2026)
von: Rezaeimanesh, Sara, et al.
Veröffentlicht: (2026)
Hallucination as Commitment Failure: Larger LLMs Misfire Despite Knowing the Answer
von: Yeom, Jewon, et al.
Veröffentlicht: (2026)
von: Yeom, Jewon, et al.
Veröffentlicht: (2026)
R-Tuning: Instructing Large Language Models to Say `I Don't Know'
von: Zhang, Hanning, et al.
Veröffentlicht: (2023)
von: Zhang, Hanning, et al.
Veröffentlicht: (2023)
Aligning What LLMs Do and Say: Towards Self-Consistent Explanations
von: Admoni, Sahar, et al.
Veröffentlicht: (2025)
von: Admoni, Sahar, et al.
Veröffentlicht: (2025)
Detecting Contextual Hallucinations in LLMs with Frequency-Aware Attention
von: Qi, Siya, et al.
Veröffentlicht: (2026)
von: Qi, Siya, et al.
Veröffentlicht: (2026)
Do LLMs Know about Hallucination? An Empirical Investigation of LLM's Hidden States
von: Duan, Hanyu, et al.
Veröffentlicht: (2024)
von: Duan, Hanyu, et al.
Veröffentlicht: (2024)
HalluShift: Measuring Distribution Shifts towards Hallucination Detection in LLMs
von: Dasgupta, Sharanya, et al.
Veröffentlicht: (2025)
von: Dasgupta, Sharanya, et al.
Veröffentlicht: (2025)
Do We Know What LLMs Don't Know? A Study of Consistency in Knowledge Probing
von: Zhao, Raoyuan, et al.
Veröffentlicht: (2025)
von: Zhao, Raoyuan, et al.
Veröffentlicht: (2025)
EpiCaR: Knowing What You Don't Know Matters for Better Reasoning in LLMs
von: Yeom, Jewon, et al.
Veröffentlicht: (2026)
von: Yeom, Jewon, et al.
Veröffentlicht: (2026)
LLMs Know What to Drop: Self-Attention Guided KV Cache Eviction for Efficient Long-Context Inference
von: Wang, Guangtao, et al.
Veröffentlicht: (2025)
von: Wang, Guangtao, et al.
Veröffentlicht: (2025)
Why Not Act on What You Know? Unleashing Safety Potential of LLMs via Self-Aware Guard Enhancement
von: Ding, Peng, et al.
Veröffentlicht: (2025)
von: Ding, Peng, et al.
Veröffentlicht: (2025)
Do LLMs Really Know What They Don't Know? Internal States Mainly Reflect Knowledge Recall Rather Than Truthfulness
von: Cheang, Chi Seng, et al.
Veröffentlicht: (2025)
von: Cheang, Chi Seng, et al.
Veröffentlicht: (2025)
Gradient-guided Attention Map Editing: Towards Efficient Contextual Hallucination Mitigation
von: Wang, Yu, et al.
Veröffentlicht: (2025)
von: Wang, Yu, et al.
Veröffentlicht: (2025)
Can AI Assistants Know What They Don't Know?
von: Cheng, Qinyuan, et al.
Veröffentlicht: (2024)
von: Cheng, Qinyuan, et al.
Veröffentlicht: (2024)
Knowing What LLMs DO NOT Know: A Simple Yet Effective Self-Detection Method
von: Zhao, Yukun, et al.
Veröffentlicht: (2023)
von: Zhao, Yukun, et al.
Veröffentlicht: (2023)
Beyond Superficial Unlearning: Sharpness-Aware Robust Erasure of Hallucinations in Multimodal LLMs
von: Fang, Xianya, et al.
Veröffentlicht: (2026)
von: Fang, Xianya, et al.
Veröffentlicht: (2026)
LLMs Know More Than They Show: On the Intrinsic Representation of LLM Hallucinations
von: Orgad, Hadas, et al.
Veröffentlicht: (2024)
von: Orgad, Hadas, et al.
Veröffentlicht: (2024)
Hallucination Detection in LLMs with Topological Divergence on Attention Graphs
von: Bazarova, Alexandra, et al.
Veröffentlicht: (2025)
von: Bazarova, Alexandra, et al.
Veröffentlicht: (2025)
LLMs on a Budget? Say HOLA
von: Siddiqui, Zohaib Hasan, et al.
Veröffentlicht: (2025)
von: Siddiqui, Zohaib Hasan, et al.
Veröffentlicht: (2025)
Fine-Tuned LLMs Know They Don't Know: A Parameter-Efficient Approach to Recovering Honesty
von: Shi, Zeyu, et al.
Veröffentlicht: (2025)
von: Shi, Zeyu, et al.
Veröffentlicht: (2025)
Decomposed Prompting Does Not Fix Knowledge Gaps, But Helps Models Say "I Don't Know"
von: Madhwal, Dhruv, et al.
Veröffentlicht: (2026)
von: Madhwal, Dhruv, et al.
Veröffentlicht: (2026)
Trust Me, I'm Wrong: LLMs Hallucinate with Certainty Despite Knowing the Answer
von: Simhi, Adi, et al.
Veröffentlicht: (2025)
von: Simhi, Adi, et al.
Veröffentlicht: (2025)
Causal Fingerprints of AI Generative Models
von: Xu, Hui, et al.
Veröffentlicht: (2025)
von: Xu, Hui, et al.
Veröffentlicht: (2025)
Shadows in the Attention: Contextual Perturbation and Representation Drift in the Dynamics of Hallucination in LLMs
von: Wei, Zeyu, et al.
Veröffentlicht: (2025)
von: Wei, Zeyu, et al.
Veröffentlicht: (2025)
DataPuzzle: Breaking Free from the Hallucinated Promise of LLMs in Data Analysis
von: Zhang, Zhengxuan, et al.
Veröffentlicht: (2025)
von: Zhang, Zhengxuan, et al.
Veröffentlicht: (2025)
KnowRL: Teaching Language Models to Know What They Know
von: Kale, Sahil, et al.
Veröffentlicht: (2025)
von: Kale, Sahil, et al.
Veröffentlicht: (2025)
Can Hallucinations Help? Boosting LLMs for Drug Discovery
von: Yuan, Shuzhou, et al.
Veröffentlicht: (2025)
von: Yuan, Shuzhou, et al.
Veröffentlicht: (2025)
LLM Ghostbusters: Surgical Hallucination Suppression via Adaptive Unlearning
von: Spracklen, Joseph, et al.
Veröffentlicht: (2026)
von: Spracklen, Joseph, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Less is More: Geometric Unlearning for LLMs with Minimal Data Disclosure
von: Tan, Chenchen, et al.
Veröffentlicht: (2026) -
When Harmless Words Harm: A New Threat to LLM Safety via Conceptual Triggers
von: Zhang, Zhaoxin, et al.
Veröffentlicht: (2025) -
Large Language Models Know What To Say But Not When To Speak
von: Umair, Muhammad, et al.
Veröffentlicht: (2024) -
You Know What I'm Saying: Jailbreak Attack via Implicit Reference
von: Wu, Tianyu, et al.
Veröffentlicht: (2024) -
LLMs Know More About Numbers than They Can Say
von: Yuchi, Fengting, et al.
Veröffentlicht: (2026)