The Curious Case of Factuality Finetuning: Models' Internal Beliefs Can Improve Factuality

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Newman, Benjamin, Ravichander, Abhilasha, Jung, Jaehun, Xin, Rui, Ivison, Hamish, Kuznetsov, Yegor, Koh, Pang Wei, Choi, Yejin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916838636519424
author Newman, Benjamin
Ravichander, Abhilasha
Jung, Jaehun
Xin, Rui
Ivison, Hamish
Kuznetsov, Yegor
Koh, Pang Wei
Choi, Yejin
author_facet Newman, Benjamin
Ravichander, Abhilasha
Jung, Jaehun
Xin, Rui
Ivison, Hamish
Kuznetsov, Yegor
Koh, Pang Wei
Choi, Yejin
contents Language models are prone to hallucination - generating text that is factually incorrect. Finetuning models on high-quality factual information can potentially reduce hallucination, but concerns remain; obtaining factual gold data can be expensive and training on correct but unfamiliar data may potentially lead to even more downstream hallucination. What data should practitioners finetune on to mitigate hallucinations in language models? In this work, we study the relationship between the factuality of finetuning data and the prevalence of hallucinations in long-form generation tasks. Counterintuitively, we find that finetuning on factual gold data is not as helpful as finetuning on model-generated data that models believe to be factual. Next, we evaluate filtering strategies applied on both factual gold data and model-generated data, and find that finetuning on model-generated data that is filtered by models' own internal judgments often leads to better overall factuality compared to other configurations: training on gold data filtered by models' judgments, training on gold data alone, or training on model-generated data that is supported by gold data. These factuality improvements transfer across three domains we study, suggesting that a models' own beliefs can provide a powerful signal for factuality.
format Preprint
id arxiv_https___arxiv_org_abs_2507_08371
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle The Curious Case of Factuality Finetuning: Models' Internal Beliefs Can Improve Factuality
Newman, Benjamin
Ravichander, Abhilasha
Jung, Jaehun
Xin, Rui
Ivison, Hamish
Kuznetsov, Yegor
Koh, Pang Wei
Choi, Yejin
Computation and Language
Language models are prone to hallucination - generating text that is factually incorrect. Finetuning models on high-quality factual information can potentially reduce hallucination, but concerns remain; obtaining factual gold data can be expensive and training on correct but unfamiliar data may potentially lead to even more downstream hallucination. What data should practitioners finetune on to mitigate hallucinations in language models? In this work, we study the relationship between the factuality of finetuning data and the prevalence of hallucinations in long-form generation tasks. Counterintuitively, we find that finetuning on factual gold data is not as helpful as finetuning on model-generated data that models believe to be factual. Next, we evaluate filtering strategies applied on both factual gold data and model-generated data, and find that finetuning on model-generated data that is filtered by models' own internal judgments often leads to better overall factuality compared to other configurations: training on gold data filtered by models' judgments, training on gold data alone, or training on model-generated data that is supported by gold data. These factuality improvements transfer across three domains we study, suggesting that a models' own beliefs can provide a powerful signal for factuality.
title The Curious Case of Factuality Finetuning: Models' Internal Beliefs Can Improve Factuality
topic Computation and Language
url https://arxiv.org/abs/2507.08371