Monitoring Decoding: Mitigating Hallucination via Evaluating the Factuality of Partial Response during Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chang, Yurui, Cao, Bochuan, Lin, Lu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911151938338816
author Chang, Yurui
Cao, Bochuan
Lin, Lu
author_facet Chang, Yurui
Cao, Bochuan
Lin, Lu
contents While large language models have demonstrated exceptional performance across a wide range of tasks, they remain susceptible to hallucinations -- generating plausible yet factually incorrect contents. Existing methods to mitigating such risk often rely on sampling multiple full-length generations, which introduces significant response latency and becomes ineffective when the model consistently produces hallucinated outputs with high confidence. To address these limitations, we introduce Monitoring Decoding (MD), a novel framework that dynamically monitors the generation process and selectively applies in-process interventions, focusing on revising crucial tokens responsible for hallucinations. Instead of waiting until completion of multiple full-length generations, we identify hallucination-prone tokens during generation using a monitor function, and further refine these tokens through a tree-based decoding strategy. This approach ensures an enhanced factual accuracy and coherence in the generated output while maintaining efficiency. Experimental results demonstrate that MD consistently outperforms self-consistency-based approaches in both effectiveness and efficiency, achieving higher factual accuracy while significantly reducing computational overhead.
format Preprint
id arxiv_https___arxiv_org_abs_2503_03106
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Monitoring Decoding: Mitigating Hallucination via Evaluating the Factuality of Partial Response during Generation
Chang, Yurui
Cao, Bochuan
Lin, Lu
Computation and Language
Machine Learning
While large language models have demonstrated exceptional performance across a wide range of tasks, they remain susceptible to hallucinations -- generating plausible yet factually incorrect contents. Existing methods to mitigating such risk often rely on sampling multiple full-length generations, which introduces significant response latency and becomes ineffective when the model consistently produces hallucinated outputs with high confidence. To address these limitations, we introduce Monitoring Decoding (MD), a novel framework that dynamically monitors the generation process and selectively applies in-process interventions, focusing on revising crucial tokens responsible for hallucinations. Instead of waiting until completion of multiple full-length generations, we identify hallucination-prone tokens during generation using a monitor function, and further refine these tokens through a tree-based decoding strategy. This approach ensures an enhanced factual accuracy and coherence in the generated output while maintaining efficiency. Experimental results demonstrate that MD consistently outperforms self-consistency-based approaches in both effectiveness and efficiency, achieving higher factual accuracy while significantly reducing computational overhead.
title Monitoring Decoding: Mitigating Hallucination via Evaluating the Factuality of Partial Response during Generation
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2503.03106