Generative Compositor for Few-Shot Visual Information Extraction

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Yang, Zhibo, Hua, Wei, Song, Sibo, Yao, Cong, Zhu, Yingying, Cheng, Wenqing, Bai, Xiang
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866916658926321664
author Yang, Zhibo
Hua, Wei
Song, Sibo
Yao, Cong
Zhu, Yingying
Cheng, Wenqing
Bai, Xiang
author_facet Yang, Zhibo
Hua, Wei
Song, Sibo
Yao, Cong
Zhu, Yingying
Cheng, Wenqing
Bai, Xiang
contents Visual Information Extraction (VIE), aiming at extracting structured information from visually rich document images, plays a pivotal role in document processing. Considering various layouts, semantic scopes, and languages, VIE encompasses an extensive range of types, potentially numbering in the thousands. However, many of these types suffer from a lack of training data, which poses significant challenges. In this paper, we propose a novel generative model, named Generative Compositor, to address the challenge of few-shot VIE. The Generative Compositor is a hybrid pointer-generator network that emulates the operations of a compositor by retrieving words from the source text and assembling them based on the provided prompts. Furthermore, three pre-training strategies are employed to enhance the model's perception of spatial context information. Besides, a prompt-aware resampler is specially designed to enable efficient matching by leveraging the entity-semantic prior contained in prompts. The introduction of the prompt-based retrieval mechanism and the pre-training strategies enable the model to acquire more effective spatial and semantic clues with limited training samples. Experiments demonstrate that the proposed method achieves highly competitive results in the full-sample training, while notably outperforms the baseline in the 1-shot, 5-shot, and 10-shot settings.
format Preprint
id arxiv_https___arxiv_org_abs_2503_16854
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Generative Compositor for Few-Shot Visual Information Extraction
Yang, Zhibo
Hua, Wei
Song, Sibo
Yao, Cong
Zhu, Yingying
Cheng, Wenqing
Bai, Xiang
Computer Vision and Pattern Recognition
Visual Information Extraction (VIE), aiming at extracting structured information from visually rich document images, plays a pivotal role in document processing. Considering various layouts, semantic scopes, and languages, VIE encompasses an extensive range of types, potentially numbering in the thousands. However, many of these types suffer from a lack of training data, which poses significant challenges. In this paper, we propose a novel generative model, named Generative Compositor, to address the challenge of few-shot VIE. The Generative Compositor is a hybrid pointer-generator network that emulates the operations of a compositor by retrieving words from the source text and assembling them based on the provided prompts. Furthermore, three pre-training strategies are employed to enhance the model's perception of spatial context information. Besides, a prompt-aware resampler is specially designed to enable efficient matching by leveraging the entity-semantic prior contained in prompts. The introduction of the prompt-based retrieval mechanism and the pre-training strategies enable the model to acquire more effective spatial and semantic clues with limited training samples. Experiments demonstrate that the proposed method achieves highly competitive results in the full-sample training, while notably outperforms the baseline in the 1-shot, 5-shot, and 10-shot settings.
title Generative Compositor for Few-Shot Visual Information Extraction
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.16854