Beyond Human Annotation: Recent Advances in Data Generation Methods for Document Intelligence

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ying, Dehao, Yu, Fengchang, Chen, Haihua, Jiang, Changjiang, Li, Yurong, Lu, Wei
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909994132176896
author Ying, Dehao
Yu, Fengchang
Chen, Haihua
Jiang, Changjiang
Li, Yurong
Lu, Wei
author_facet Ying, Dehao
Yu, Fengchang
Chen, Haihua
Jiang, Changjiang
Li, Yurong
Lu, Wei
contents The advancement of Document Intelligence (DI) demands large-scale, high-quality training data, yet manual annotation remains a critical bottleneck. While data generation methods are evolving rapidly, existing surveys are constrained by fragmented focuses on single modalities or specific tasks, lacking a unified perspective aligned with real-world workflows. To fill this gap, this survey establishes the first comprehensive technical map for data generation in DI. Data generation is redefined as supervisory signal production, and a novel taxonomy is introduced based on the "availability of data and labels." This framework organizes methodologies into four resource-centric paradigms: Data Augmentation, Data Generation from Scratch, Automated Data Annotation, and Self-Supervised Signal Construction. Furthermore, a multi-level evaluation framework is established to integrate intrinsic quality and extrinsic utility, compiling performance gains across diverse DI benchmarks. Guided by this unified structure, the methodological landscape is dissected to reveal critical challenges such as fidelity gaps and frontiers including co-evolutionary ecosystems. Ultimately, by systematizing this fragmented field, data generation is positioned as the central engine for next-generation DI.
format Preprint
id arxiv_https___arxiv_org_abs_2601_12318
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Beyond Human Annotation: Recent Advances in Data Generation Methods for Document Intelligence
Ying, Dehao
Yu, Fengchang
Chen, Haihua
Jiang, Changjiang
Li, Yurong
Lu, Wei
Artificial Intelligence
The advancement of Document Intelligence (DI) demands large-scale, high-quality training data, yet manual annotation remains a critical bottleneck. While data generation methods are evolving rapidly, existing surveys are constrained by fragmented focuses on single modalities or specific tasks, lacking a unified perspective aligned with real-world workflows. To fill this gap, this survey establishes the first comprehensive technical map for data generation in DI. Data generation is redefined as supervisory signal production, and a novel taxonomy is introduced based on the "availability of data and labels." This framework organizes methodologies into four resource-centric paradigms: Data Augmentation, Data Generation from Scratch, Automated Data Annotation, and Self-Supervised Signal Construction. Furthermore, a multi-level evaluation framework is established to integrate intrinsic quality and extrinsic utility, compiling performance gains across diverse DI benchmarks. Guided by this unified structure, the methodological landscape is dissected to reveal critical challenges such as fidelity gaps and frontiers including co-evolutionary ecosystems. Ultimately, by systematizing this fragmented field, data generation is positioned as the central engine for next-generation DI.
title Beyond Human Annotation: Recent Advances in Data Generation Methods for Document Intelligence
topic Artificial Intelligence
url https://arxiv.org/abs/2601.12318