Honest AI: Fine-Tuning "Small" Language Models to Say "I Don't Know", and Reducing Hallucination in RAG
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Chen, Xinxi, Wang, Li, Wu, Wei, Tang, Qi, Liu, Yiyao |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
What Models Know, How Well They Know It: Knowledge-Weighted Fine-Tuning for Learning When to Say "I Don't Know"
par: Lee, Joosung, et autres
Publié: (2026)
par: Lee, Joosung, et autres
Publié: (2026)
R-Tuning: Instructing Large Language Models to Say `I Don't Know'
par: Zhang, Hanning, et autres
Publié: (2023)
par: Zhang, Hanning, et autres
Publié: (2023)
Can AI Assistants Know What They Don't Know?
par: Cheng, Qinyuan, et autres
Publié: (2024)
par: Cheng, Qinyuan, et autres
Publié: (2024)
Do Retrieval Augmented Language Models Know When They Don't Know?
par: Zhou, Youchao, et autres
Publié: (2025)
par: Zhou, Youchao, et autres
Publié: (2025)
Large Language Models Must Be Taught to Know What They Don't Know
par: Kapoor, Sanyam, et autres
Publié: (2024)
par: Kapoor, Sanyam, et autres
Publié: (2024)
Reasoning Models Don't Always Say What They Think
par: Chen, Yanda, et autres
Publié: (2025)
par: Chen, Yanda, et autres
Publié: (2025)
Reasoning about Uncertainty: Do Reasoning Models Know When They Don't Know?
par: Mei, Zhiting, et autres
Publié: (2025)
par: Mei, Zhiting, et autres
Publié: (2025)
Knowing You Don't Know: Learning When to Continue Search in Multi-round RAG through Self-Practicing
par: Yang, Diji, et autres
Publié: (2025)
par: Yang, Diji, et autres
Publié: (2025)
Decomposed Prompting Does Not Fix Knowledge Gaps, But Helps Models Say "I Don't Know"
par: Madhwal, Dhruv, et autres
Publié: (2026)
par: Madhwal, Dhruv, et autres
Publié: (2026)
Fine-Tuned LLMs Know They Don't Know: A Parameter-Efficient Approach to Recovering Honesty
par: Shi, Zeyu, et autres
Publié: (2025)
par: Shi, Zeyu, et autres
Publié: (2025)
What Language Models Know But Don't Say: Non-Generative Prior Extraction for Generalization
par: Rezaeimanesh, Sara, et autres
Publié: (2026)
par: Rezaeimanesh, Sara, et autres
Publié: (2026)
HonestLLM: Toward an Honest and Helpful Large Language Model
par: Gao, Chujie, et autres
Publié: (2024)
par: Gao, Chujie, et autres
Publié: (2024)
Coarse-to-Fine Highlighting: Reducing Knowledge Hallucination in Large Language Models
par: Lv, Qitan, et autres
Publié: (2024)
par: Lv, Qitan, et autres
Publié: (2024)
Don't Fine-Tune, Decode: Syntax Error-Free Tool Use via Constrained Decoding
par: Zhang, Kexun, et autres
Publié: (2023)
par: Zhang, Kexun, et autres
Publié: (2023)
Why Models Know But Don't Say: Chain-of-Thought Faithfulness Divergence Between Thinking Tokens and Answers in Open-Weight Reasoning Models
par: Young, Richard J.
Publié: (2026)
par: Young, Richard J.
Publié: (2026)
Fine-Tuning Language Models to Know What They Know
par: Park, Sangjun, et autres
Publié: (2026)
par: Park, Sangjun, et autres
Publié: (2026)
KnowTuning: Knowledge-aware Fine-tuning for Large Language Models
par: Lyu, Yougang, et autres
Publié: (2024)
par: Lyu, Yougang, et autres
Publié: (2024)
Mini-Giants: "Small" Language Models and Open Source Win-Win
par: Zhou, Zhengping, et autres
Publié: (2023)
par: Zhou, Zhengping, et autres
Publié: (2023)
ToolBeHonest: A Multi-level Hallucination Diagnostic Benchmark for Tool-Augmented Large Language Models
par: Zhang, Yuxiang, et autres
Publié: (2024)
par: Zhang, Yuxiang, et autres
Publié: (2024)
Language Models Don't Learn the Physical Manifestation of Language
par: Lee, Bruce W., et autres
Publié: (2024)
par: Lee, Bruce W., et autres
Publié: (2024)
When Robots Should Say "I Don't Know": Benchmarking Abstention in Embodied Question Answering
par: Wu, Tao, et autres
Publié: (2025)
par: Wu, Tao, et autres
Publié: (2025)
The Accuracy Paradox in RLHF: When Better Reward Models Don't Yield Better Language Models
par: Chen, Yanjun, et autres
Publié: (2024)
par: Chen, Yanjun, et autres
Publié: (2024)
Fine-Tuned Large Language Models for Logical Translation: Reducing Hallucinations with Lang2Logic
par: Pan, Muyu, et autres
Publié: (2025)
par: Pan, Muyu, et autres
Publié: (2025)
Don't Believe Everything You Read: Enhancing Summarization Interpretability through Automatic Identification of Hallucinations in Large Language Models
par: Vakharia, Priyesh, et autres
Publié: (2023)
par: Vakharia, Priyesh, et autres
Publié: (2023)
Don't Let It Hallucinate: Premise Verification via Retrieval-Augmented Logical Reasoning
par: Qin, Yuehan, et autres
Publié: (2025)
par: Qin, Yuehan, et autres
Publié: (2025)
Chip-Tuning: Classify Before Language Models Say
par: Zhu, Fangwei, et autres
Publié: (2024)
par: Zhu, Fangwei, et autres
Publié: (2024)
Noise Augmented Fine Tuning for Mitigating Hallucinations in Large Language Models
par: Khadangi, Afshin, et autres
Publié: (2025)
par: Khadangi, Afshin, et autres
Publié: (2025)
I Don't Know: Explicit Modeling of Uncertainty with an [IDK] Token
par: Cohen, Roi, et autres
Publié: (2024)
par: Cohen, Roi, et autres
Publié: (2024)
Do I Know This Entity? Knowledge Awareness and Hallucinations in Language Models
par: Ferrando, Javier, et autres
Publié: (2024)
par: Ferrando, Javier, et autres
Publié: (2024)
Know3-RAG: A Knowledge-aware RAG Framework with Adaptive Retrieval, Generation, and Filtering
par: Liu, Xukai, et autres
Publié: (2025)
par: Liu, Xukai, et autres
Publié: (2025)
Measuring What VLMs Don't Say: Validation Metrics Hide Clinical Terminology Erasure in Radiology Report Generation
par: Parikh, Aditya, et autres
Publié: (2026)
par: Parikh, Aditya, et autres
Publié: (2026)
Don't Say No: Jailbreaking LLM by Suppressing Refusal
par: Zhou, Yukai, et autres
Publié: (2024)
par: Zhou, Yukai, et autres
Publié: (2024)
Can LLMs Ground when they (Don't) Know: A Study on Direct and Loaded Political Questions
par: Lachenmaier, Clara, et autres
Publié: (2025)
par: Lachenmaier, Clara, et autres
Publié: (2025)
Removal of Hallucination on Hallucination: Debate-Augmented RAG
par: Hu, Wentao, et autres
Publié: (2025)
par: Hu, Wentao, et autres
Publié: (2025)
Don't Retrieve, Navigate: Distilling Enterprise Knowledge into Navigable Agent Skills for QA and RAG
par: Sun, Yiqun, et autres
Publié: (2026)
par: Sun, Yiqun, et autres
Publié: (2026)
Don't Forget to Connect! Improving RAG with Graph-based Reranking
par: Dong, Jialin, et autres
Publié: (2024)
par: Dong, Jialin, et autres
Publié: (2024)
Don't Pay Attention
par: Hammoud, Mohammad, et autres
Publié: (2025)
par: Hammoud, Mohammad, et autres
Publié: (2025)
LLMs Don't Know Their Own Decision Boundaries: The Unreliability of Self-Generated Counterfactual Explanations
par: Mayne, Harry, et autres
Publié: (2025)
par: Mayne, Harry, et autres
Publié: (2025)
BeHonest: Benchmarking Honesty in Large Language Models
par: Chern, Steffi, et autres
Publié: (2024)
par: Chern, Steffi, et autres
Publié: (2024)
Fine-Tune, Don't Prompt, Your Language Model to Identify Biased Language in Clinical Notes
par: Landi, Isotta, et autres
Publié: (2026)
par: Landi, Isotta, et autres
Publié: (2026)
Documents similaires
-
What Models Know, How Well They Know It: Knowledge-Weighted Fine-Tuning for Learning When to Say "I Don't Know"
par: Lee, Joosung, et autres
Publié: (2026) -
R-Tuning: Instructing Large Language Models to Say `I Don't Know'
par: Zhang, Hanning, et autres
Publié: (2023) -
Can AI Assistants Know What They Don't Know?
par: Cheng, Qinyuan, et autres
Publié: (2024) -
Do Retrieval Augmented Language Models Know When They Don't Know?
par: Zhou, Youchao, et autres
Publié: (2025) -
Large Language Models Must Be Taught to Know What They Don't Know
par: Kapoor, Sanyam, et autres
Publié: (2024)