FAITH: Factuality Alignment through Integrating Trustworthiness and Honestness

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dong, Xiaoning, Wu, Chengyan, Wen, Yajie, Chen, Yu, Xue, Yun, Zhang, Jing, Xu, Wei, Ma, Bolei
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914465042137088
author Dong, Xiaoning
Wu, Chengyan
Wen, Yajie
Chen, Yu
Xue, Yun
Zhang, Jing
Xu, Wei
Ma, Bolei
author_facet Dong, Xiaoning
Wu, Chengyan
Wen, Yajie
Chen, Yu
Xue, Yun
Zhang, Jing
Xu, Wei
Ma, Bolei
contents Large Language Models (LLMs) can generate factually inaccurate content even if they have corresponding knowledge, which critically undermines their reliability. Existing approaches attempt to mitigate this by incorporating uncertainty in QA prompt during training, but these numerical scores lack the semantic richness for LLM to properly understand its internal states of trustworthiness and honestness, leading to insufficient factuality alignment. We introduce FAITH (Factuality Alignment through Integrating Trustworthiness and Honestness), a post-training framework for factuality alignment that integrates natural-language uncertainty signals with external knowledge. Specifically, we augment training datasets by computing confidence scores and semantic entropy from LLM outputs and mapping them into a knowledge state quadrant that describes the model's internal knowledge possession (trustworthiness) and answering behaviors (honestness) in natural language. Based on this enhanced data, we design a reward function that considers both correctness and uncertainty signals, and fine-tune the LLM using the Proximal Policy Optimization (PPO) algorithm. To further mitigate weakly grounded responses, we design a retrieval-augmented module that retrieves relevant external passages, improving the consistency between internal and external knowledge representations. Extensive experiments on four knowledge-intensive benchmarks demonstrate that FAITH enhances the factual accuracy and truthfulness of LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2604_10189
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle FAITH: Factuality Alignment through Integrating Trustworthiness and Honestness
Dong, Xiaoning
Wu, Chengyan
Wen, Yajie
Chen, Yu
Xue, Yun
Zhang, Jing
Xu, Wei
Ma, Bolei
Computation and Language
Large Language Models (LLMs) can generate factually inaccurate content even if they have corresponding knowledge, which critically undermines their reliability. Existing approaches attempt to mitigate this by incorporating uncertainty in QA prompt during training, but these numerical scores lack the semantic richness for LLM to properly understand its internal states of trustworthiness and honestness, leading to insufficient factuality alignment. We introduce FAITH (Factuality Alignment through Integrating Trustworthiness and Honestness), a post-training framework for factuality alignment that integrates natural-language uncertainty signals with external knowledge. Specifically, we augment training datasets by computing confidence scores and semantic entropy from LLM outputs and mapping them into a knowledge state quadrant that describes the model's internal knowledge possession (trustworthiness) and answering behaviors (honestness) in natural language. Based on this enhanced data, we design a reward function that considers both correctness and uncertainty signals, and fine-tune the LLM using the Proximal Policy Optimization (PPO) algorithm. To further mitigate weakly grounded responses, we design a retrieval-augmented module that retrieves relevant external passages, improving the consistency between internal and external knowledge representations. Extensive experiments on four knowledge-intensive benchmarks demonstrate that FAITH enhances the factual accuracy and truthfulness of LLMs.
title FAITH: Factuality Alignment through Integrating Trustworthiness and Honestness
topic Computation and Language
url https://arxiv.org/abs/2604.10189