On Finding Inconsistencies in Documents

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lovering, Charles J., Ebner, Seth, Smock, Brandon, Krumdick, Michael, Rabbani, Saad, Muhammad, Ahmed, Reddy, Varshini, Tanner, Chris
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914212934057984
author Lovering, Charles J.
Ebner, Seth
Smock, Brandon
Krumdick, Michael
Rabbani, Saad
Muhammad, Ahmed
Reddy, Varshini
Tanner, Chris
author_facet Lovering, Charles J.
Ebner, Seth
Smock, Brandon
Krumdick, Michael
Rabbani, Saad
Muhammad, Ahmed
Reddy, Varshini
Tanner, Chris
contents Professionals in academia, law, and finance audit their documents because inconsistencies can result in monetary, reputational, and scientific costs. Language models (LMs) have the potential to dramatically speed up this auditing process. To understand their abilities, we introduce a benchmark, FIND (Finding INconsistencies in Documents), where each example is a document with an inconsistency inserted manually by a domain expert. Despite the documents being long, technical, and complex, the best-performing model (gpt-5) recovered 64% of the inserted inconsistencies. Surprisingly, gpt-5 also found undiscovered inconsistencies present in the original documents. For example, on 50 arXiv papers, we judged 136 out of 196 of the model's suggestions to be legitimate inconsistencies missed by the original authors. However, despite these findings, even the best models miss almost half of the inconsistencies in FIND, demonstrating that inconsistency detection is still a challenging task.
format Preprint
id arxiv_https___arxiv_org_abs_2512_18601
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle On Finding Inconsistencies in Documents
Lovering, Charles J.
Ebner, Seth
Smock, Brandon
Krumdick, Michael
Rabbani, Saad
Muhammad, Ahmed
Reddy, Varshini
Tanner, Chris
Computation and Language
Professionals in academia, law, and finance audit their documents because inconsistencies can result in monetary, reputational, and scientific costs. Language models (LMs) have the potential to dramatically speed up this auditing process. To understand their abilities, we introduce a benchmark, FIND (Finding INconsistencies in Documents), where each example is a document with an inconsistency inserted manually by a domain expert. Despite the documents being long, technical, and complex, the best-performing model (gpt-5) recovered 64% of the inserted inconsistencies. Surprisingly, gpt-5 also found undiscovered inconsistencies present in the original documents. For example, on 50 arXiv papers, we judged 136 out of 196 of the model's suggestions to be legitimate inconsistencies missed by the original authors. However, despite these findings, even the best models miss almost half of the inconsistencies in FIND, demonstrating that inconsistency detection is still a challenging task.
title On Finding Inconsistencies in Documents
topic Computation and Language
url https://arxiv.org/abs/2512.18601