TruthRL: Incentivizing Truthful LLMs via Reinforcement Learning

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Wei, Zhepei, Yang, Xiao, Sun, Kai, Wang, Jiaqi, Shao, Rulin, Chen, Sean, Kachuee, Mohammad, Gollapudi, Teja, Liao, Tony, Scheffer, Nicolas, Wanga, Rakesh, Kumar, Anuj, Meng, Yu, Yih, Wen-tau, Dong, Xin Luna
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916979316621312
author Wei, Zhepei
Yang, Xiao
Sun, Kai
Wang, Jiaqi
Shao, Rulin
Chen, Sean
Kachuee, Mohammad
Gollapudi, Teja
Liao, Tony
Scheffer, Nicolas
Wanga, Rakesh
Kumar, Anuj
Meng, Yu
Yih, Wen-tau
Dong, Xin Luna
author_facet Wei, Zhepei
Yang, Xiao
Sun, Kai
Wang, Jiaqi
Shao, Rulin
Chen, Sean
Kachuee, Mohammad
Gollapudi, Teja
Liao, Tony
Scheffer, Nicolas
Wanga, Rakesh
Kumar, Anuj
Meng, Yu
Yih, Wen-tau
Dong, Xin Luna
contents While large language models (LLMs) have demonstrated strong performance on factoid question answering, they are still prone to hallucination and untruthful responses, particularly when tasks demand information outside their parametric knowledge. Indeed, truthfulness requires more than accuracy -- models must also recognize uncertainty and abstain when unsure to avoid hallucinations. This presents a fundamental challenge for existing methods: approaches that optimize for accuracy often amplify hallucinations, while those that encourage abstention can become overly conservative, sacrificing correct answers. Both extremes ultimately compromise truthfulness. In this work, we present TruthRL, a general reinforcement learning (RL) framework that directly optimizes the truthfulness of LLMs. Specifically, we implement TruthRL using GRPO with a simple yet effective ternary reward that distinguishes correct answers, hallucinations, and abstentions. It incentivizes models to reduce hallucinations not only by providing correct responses, but also by enabling abstention when uncertain, thereby improving truthfulness. Extensive experiments across four knowledge-intensive benchmarks show that, compared to vanilla RL, TruthRL significantly reduces hallucinations by 28.9% and improves truthfulness by 21.1%, with consistent gains across various backbone models (e.g., Qwen, Llama) under both retrieval and non-retrieval setups. In-depth ablation study demonstrates that vanilla accuracy-driven methods, such as supervised fine-tuning or RL with a binary reward, struggle to balance factual correctness and uncertainty. In contrast, our proposed truthfulness-driven TruthRL achieves strong performance in both accuracy and truthfulness, underscoring the importance of learning objective design for developing truthful LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2509_25760
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle TruthRL: Incentivizing Truthful LLMs via Reinforcement Learning
Wei, Zhepei
Yang, Xiao
Sun, Kai
Wang, Jiaqi
Shao, Rulin
Chen, Sean
Kachuee, Mohammad
Gollapudi, Teja
Liao, Tony
Scheffer, Nicolas
Wanga, Rakesh
Kumar, Anuj
Meng, Yu
Yih, Wen-tau
Dong, Xin Luna
Computation and Language
Artificial Intelligence
Machine Learning
While large language models (LLMs) have demonstrated strong performance on factoid question answering, they are still prone to hallucination and untruthful responses, particularly when tasks demand information outside their parametric knowledge. Indeed, truthfulness requires more than accuracy -- models must also recognize uncertainty and abstain when unsure to avoid hallucinations. This presents a fundamental challenge for existing methods: approaches that optimize for accuracy often amplify hallucinations, while those that encourage abstention can become overly conservative, sacrificing correct answers. Both extremes ultimately compromise truthfulness. In this work, we present TruthRL, a general reinforcement learning (RL) framework that directly optimizes the truthfulness of LLMs. Specifically, we implement TruthRL using GRPO with a simple yet effective ternary reward that distinguishes correct answers, hallucinations, and abstentions. It incentivizes models to reduce hallucinations not only by providing correct responses, but also by enabling abstention when uncertain, thereby improving truthfulness. Extensive experiments across four knowledge-intensive benchmarks show that, compared to vanilla RL, TruthRL significantly reduces hallucinations by 28.9% and improves truthfulness by 21.1%, with consistent gains across various backbone models (e.g., Qwen, Llama) under both retrieval and non-retrieval setups. In-depth ablation study demonstrates that vanilla accuracy-driven methods, such as supervised fine-tuning or RL with a binary reward, struggle to balance factual correctness and uncertainty. In contrast, our proposed truthfulness-driven TruthRL achieves strong performance in both accuracy and truthfulness, underscoring the importance of learning objective design for developing truthful LLMs.
title TruthRL: Incentivizing Truthful LLMs via Reinforcement Learning
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2509.25760