Balancing Truthfulness and Informativeness with Uncertainty-Aware Instruction Fine-Tuning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Tianyi, Ni, Jingwei, Hooi, Bryan, Zhang, Jiaheng, Ash, Elliott, Ng, See-Kiong, Sachan, Mrinmaya, Leippold, Markus
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909659534721024
author Wu, Tianyi
Ni, Jingwei
Hooi, Bryan
Zhang, Jiaheng
Ash, Elliott
Ng, See-Kiong
Sachan, Mrinmaya
Leippold, Markus
author_facet Wu, Tianyi
Ni, Jingwei
Hooi, Bryan
Zhang, Jiaheng
Ash, Elliott
Ng, See-Kiong
Sachan, Mrinmaya
Leippold, Markus
contents Instruction fine-tuning (IFT) can increase the informativeness of large language models (LLMs), but may reduce their truthfulness. This trade-off arises because IFT steers LLMs to generate responses containing long-tail knowledge that was not well covered during pre-training. As a result, models become more informative but less accurate when generalizing to unseen tasks. In this paper, we empirically demonstrate how unfamiliar knowledge in IFT datasets can negatively affect the truthfulness of LLMs, and we introduce two new IFT paradigms, $UNIT_{cut}$ and $UNIT_{ref}$, to address this issue. $UNIT_{cut}$ identifies and removes unfamiliar knowledge from IFT datasets to mitigate its impact on model truthfulness, whereas $UNIT_{ref}$ trains LLMs to recognize their uncertainty and explicitly indicate it at the end of their responses. Our experiments show that $UNIT_{cut}$ substantially improves LLM truthfulness, while $UNIT_{ref}$ maintains high informativeness and reduces hallucinations by distinguishing between confident and uncertain statements.
format Preprint
id arxiv_https___arxiv_org_abs_2502_11962
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Balancing Truthfulness and Informativeness with Uncertainty-Aware Instruction Fine-Tuning
Wu, Tianyi
Ni, Jingwei
Hooi, Bryan
Zhang, Jiaheng
Ash, Elliott
Ng, See-Kiong
Sachan, Mrinmaya
Leippold, Markus
Computation and Language
Artificial Intelligence
Instruction fine-tuning (IFT) can increase the informativeness of large language models (LLMs), but may reduce their truthfulness. This trade-off arises because IFT steers LLMs to generate responses containing long-tail knowledge that was not well covered during pre-training. As a result, models become more informative but less accurate when generalizing to unseen tasks. In this paper, we empirically demonstrate how unfamiliar knowledge in IFT datasets can negatively affect the truthfulness of LLMs, and we introduce two new IFT paradigms, $UNIT_{cut}$ and $UNIT_{ref}$, to address this issue. $UNIT_{cut}$ identifies and removes unfamiliar knowledge from IFT datasets to mitigate its impact on model truthfulness, whereas $UNIT_{ref}$ trains LLMs to recognize their uncertainty and explicitly indicate it at the end of their responses. Our experiments show that $UNIT_{cut}$ substantially improves LLM truthfulness, while $UNIT_{ref}$ maintains high informativeness and reduces hallucinations by distinguishing between confident and uncertain statements.
title Balancing Truthfulness and Informativeness with Uncertainty-Aware Instruction Fine-Tuning
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2502.11962