Improved Evidence Extraction and Metrics for Document Inconsistency Detection with LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tan, Nelvin, Zhang, Yaowen, Cheung, James Asikin, Liu, Fusheng, Shih, Yu-Ching, Yang, Dong
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911575520051200
author Tan, Nelvin
Zhang, Yaowen
Cheung, James Asikin
Liu, Fusheng
Shih, Yu-Ching
Yang, Dong
author_facet Tan, Nelvin
Zhang, Yaowen
Cheung, James Asikin
Liu, Fusheng
Shih, Yu-Ching
Yang, Dong
contents Large language models (LLMs) are becoming useful in many domains due to their impressive abilities that arise from large training datasets and large model sizes. However, research on LLM-based approaches to document inconsistency detection is relatively limited. We address this gap by investigating evidence extraction capabilties of LLMs for document inconsistency detection. To this end, we introduce new comprehensive evidence-extraction metrics and a redact-and-retry framework with constrained filtering that substantially improves evidence extraction performance over other prompting methods. We support our approach with strong experimental results and release a new semi-synthetic dataset for evaluating evidence extraction.
format Preprint
id arxiv_https___arxiv_org_abs_2601_02627
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Improved Evidence Extraction and Metrics for Document Inconsistency Detection with LLMs
Tan, Nelvin
Zhang, Yaowen
Cheung, James Asikin
Liu, Fusheng
Shih, Yu-Ching
Yang, Dong
Computation and Language
Artificial Intelligence
Large language models (LLMs) are becoming useful in many domains due to their impressive abilities that arise from large training datasets and large model sizes. However, research on LLM-based approaches to document inconsistency detection is relatively limited. We address this gap by investigating evidence extraction capabilties of LLMs for document inconsistency detection. To this end, we introduce new comprehensive evidence-extraction metrics and a redact-and-retry framework with constrained filtering that substantially improves evidence extraction performance over other prompting methods. We support our approach with strong experimental results and release a new semi-synthetic dataset for evaluating evidence extraction.
title Improved Evidence Extraction and Metrics for Document Inconsistency Detection with LLMs
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2601.02627