Learning to Reason for Hallucination Span Detection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Su, Hsuan, Hu, Ting-Yao, Koppula, Hema Swetha, Krishna, Kundan, Pouransari, Hadi, Hsieh, Cheng-Yu, Koc, Cem, Cheng, Joseph Yitan, Tuzel, Oncel, Vemulapalli, Raviteja
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912637525164032
author Su, Hsuan
Hu, Ting-Yao
Koppula, Hema Swetha
Krishna, Kundan
Pouransari, Hadi
Hsieh, Cheng-Yu
Koc, Cem
Cheng, Joseph Yitan
Tuzel, Oncel
Vemulapalli, Raviteja
author_facet Su, Hsuan
Hu, Ting-Yao
Koppula, Hema Swetha
Krishna, Kundan
Pouransari, Hadi
Hsieh, Cheng-Yu
Koc, Cem
Cheng, Joseph Yitan
Tuzel, Oncel
Vemulapalli, Raviteja
contents Large language models (LLMs) often generate hallucinations -- unsupported content that undermines reliability. While most prior works frame hallucination detection as a binary task, many real-world applications require identifying hallucinated spans, which is a multi-step decision making process. This naturally raises the question of whether explicit reasoning can help the complex task of detecting hallucination spans. To answer this question, we first evaluate pretrained models with and without Chain-of-Thought (CoT) reasoning, and show that CoT reasoning has the potential to generate at least one correct answer when sampled multiple times. Motivated by this, we propose RL4HS, a reinforcement learning framework that incentivizes reasoning with a span-level reward function. RL4HS builds on Group Relative Policy Optimization and introduces Class-Aware Policy Optimization to mitigate reward imbalance issue. Experiments on the RAGTruth benchmark (summarization, question answering, data-to-text) show that RL4HS surpasses pretrained reasoning models and supervised fine-tuning, demonstrating the necessity of reinforcement learning with span-level rewards for detecting hallucination spans.
format Preprint
id arxiv_https___arxiv_org_abs_2510_02173
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Learning to Reason for Hallucination Span Detection
Su, Hsuan
Hu, Ting-Yao
Koppula, Hema Swetha
Krishna, Kundan
Pouransari, Hadi
Hsieh, Cheng-Yu
Koc, Cem
Cheng, Joseph Yitan
Tuzel, Oncel
Vemulapalli, Raviteja
Computation and Language
Artificial Intelligence
Machine Learning
Large language models (LLMs) often generate hallucinations -- unsupported content that undermines reliability. While most prior works frame hallucination detection as a binary task, many real-world applications require identifying hallucinated spans, which is a multi-step decision making process. This naturally raises the question of whether explicit reasoning can help the complex task of detecting hallucination spans. To answer this question, we first evaluate pretrained models with and without Chain-of-Thought (CoT) reasoning, and show that CoT reasoning has the potential to generate at least one correct answer when sampled multiple times. Motivated by this, we propose RL4HS, a reinforcement learning framework that incentivizes reasoning with a span-level reward function. RL4HS builds on Group Relative Policy Optimization and introduces Class-Aware Policy Optimization to mitigate reward imbalance issue. Experiments on the RAGTruth benchmark (summarization, question answering, data-to-text) show that RL4HS surpasses pretrained reasoning models and supervised fine-tuning, demonstrating the necessity of reinforcement learning with span-level rewards for detecting hallucination spans.
title Learning to Reason for Hallucination Span Detection
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2510.02173