GPT4o-Receipt: A Dataset and Human Study for AI-Generated Document Forensics

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Yan, Ren, Simiao, Raj, Ankit, Wei, En, Ng, Dennis, Shen, Alex, Xue, Jiayu, Zhang, Yuxin, Marotta, Evelyn
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912981209579520
author Zhang, Yan
Ren, Simiao
Raj, Ankit
Wei, En
Ng, Dennis
Shen, Alex
Xue, Jiayu
Zhang, Yuxin
Marotta, Evelyn
author_facet Zhang, Yan
Ren, Simiao
Raj, Ankit
Wei, En
Ng, Dennis
Shen, Alex
Xue, Jiayu
Zhang, Yuxin
Marotta, Evelyn
contents Can humans detect AI-generated financial documents better than machines? We present GPT4o-Receipt, a benchmark of 1,235 receipt images pairing GPT-4o-generated receipts with authentic ones from established datasets, evaluated by five state-of-the-art multimodal LLMs and a 30-annotator crowdsourced perceptual study. Our findings reveal a striking paradox: humans are better at seeing AI artifacts, yet worse at detecting AI documents. Human annotators exhibit the largest visual discrimination gap of any evaluator, yet their binary detection F1 falls well below Claude Sonnet 4 and below Gemini 2.5 Flash. This paradox resolves once the mechanism is understood: the dominant forensic signals in AI-generated receipts are arithmetic errors -- invisible to visual inspection but systematically verifiable by LLMs. Humans cannot perceive that a subtotal is incorrect; LLMs verify it in milliseconds. Beyond the human--LLM comparison, our five-model evaluation reveals dramatic performance disparities and calibration differences that render simple accuracy metrics insufficient for detector selection. GPT4o-Receipt, the evaluation framework, and all results are released publicly to support future research in AI document forensics.
format Preprint
id arxiv_https___arxiv_org_abs_2603_11442
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle GPT4o-Receipt: A Dataset and Human Study for AI-Generated Document Forensics
Zhang, Yan
Ren, Simiao
Raj, Ankit
Wei, En
Ng, Dennis
Shen, Alex
Xue, Jiayu
Zhang, Yuxin
Marotta, Evelyn
Artificial Intelligence
Computer Vision and Pattern Recognition
I.4.9; I.2.10
Can humans detect AI-generated financial documents better than machines? We present GPT4o-Receipt, a benchmark of 1,235 receipt images pairing GPT-4o-generated receipts with authentic ones from established datasets, evaluated by five state-of-the-art multimodal LLMs and a 30-annotator crowdsourced perceptual study. Our findings reveal a striking paradox: humans are better at seeing AI artifacts, yet worse at detecting AI documents. Human annotators exhibit the largest visual discrimination gap of any evaluator, yet their binary detection F1 falls well below Claude Sonnet 4 and below Gemini 2.5 Flash. This paradox resolves once the mechanism is understood: the dominant forensic signals in AI-generated receipts are arithmetic errors -- invisible to visual inspection but systematically verifiable by LLMs. Humans cannot perceive that a subtotal is incorrect; LLMs verify it in milliseconds. Beyond the human--LLM comparison, our five-model evaluation reveals dramatic performance disparities and calibration differences that render simple accuracy metrics insufficient for detector selection. GPT4o-Receipt, the evaluation framework, and all results are released publicly to support future research in AI document forensics.
title GPT4o-Receipt: A Dataset and Human Study for AI-Generated Document Forensics
topic Artificial Intelligence
Computer Vision and Pattern Recognition
I.4.9; I.2.10
url https://arxiv.org/abs/2603.11442