POINTS-Reader: Distillation-Free Adaptation of Vision-Language Models for Document Conversion

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Yuan, Zhao, Zhongyin, Tian, Le, Wang, Haicheng, Ye, Xubing, You, Yangxiu, Yu, Zilin, Wu, Chuhan, Zhou, Xiao, Yu, Yang, Zhou, Jie
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911132154855424
author Liu, Yuan
Zhao, Zhongyin
Tian, Le
Wang, Haicheng
Ye, Xubing
You, Yangxiu
Yu, Zilin
Wu, Chuhan
Zhou, Xiao
Yu, Yang
Zhou, Jie
author_facet Liu, Yuan
Zhao, Zhongyin
Tian, Le
Wang, Haicheng
Ye, Xubing
You, Yangxiu
Yu, Zilin
Wu, Chuhan
Zhou, Xiao
Yu, Yang
Zhou, Jie
contents High-quality labeled data is essential for training accurate document conversion models, particularly in domains with complex formats such as tables, formulas, and multi-column text. However, manual annotation is both costly and time-consuming, while automatic labeling using existing models often lacks accuracy in handling such challenging scenarios. Consequently, training student models by distilling outputs from teacher models can significantly limit their performance in real-world applications. In this paper, we propose a fully automated, distillation-free framework comprising two stages for constructing high-quality document extraction datasets and models capable of handling diverse document formats and layouts. In the first stage, we introduce a method for generating large-scale, diverse synthetic data, which enables a model to extract key elements in a unified format with strong initial performance. In the second stage, we present a self-improvement approach that further adapts the model, initially trained on synthetic data, to real-world documents. Specifically, we first use the fine-tuned model to annotate real documents, then apply a suite of filtering strategies to verify annotation quality, and finally retrain the model on the verified dataset. By iteratively repeating this process, we progressively enhance both the model's conversion capabilities and the quality of the generated data. We train a public POINTS-1.5 model to obtain POINTS-Reader, which surpasses many existing public and proprietary models of comparable or larger size. Our model is available at https://github.com/Tencent/POINTS-Reader.
format Preprint
id arxiv_https___arxiv_org_abs_2509_01215
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle POINTS-Reader: Distillation-Free Adaptation of Vision-Language Models for Document Conversion
Liu, Yuan
Zhao, Zhongyin
Tian, Le
Wang, Haicheng
Ye, Xubing
You, Yangxiu
Yu, Zilin
Wu, Chuhan
Zhou, Xiao
Yu, Yang
Zhou, Jie
Computer Vision and Pattern Recognition
High-quality labeled data is essential for training accurate document conversion models, particularly in domains with complex formats such as tables, formulas, and multi-column text. However, manual annotation is both costly and time-consuming, while automatic labeling using existing models often lacks accuracy in handling such challenging scenarios. Consequently, training student models by distilling outputs from teacher models can significantly limit their performance in real-world applications. In this paper, we propose a fully automated, distillation-free framework comprising two stages for constructing high-quality document extraction datasets and models capable of handling diverse document formats and layouts. In the first stage, we introduce a method for generating large-scale, diverse synthetic data, which enables a model to extract key elements in a unified format with strong initial performance. In the second stage, we present a self-improvement approach that further adapts the model, initially trained on synthetic data, to real-world documents. Specifically, we first use the fine-tuned model to annotate real documents, then apply a suite of filtering strategies to verify annotation quality, and finally retrain the model on the verified dataset. By iteratively repeating this process, we progressively enhance both the model's conversion capabilities and the quality of the generated data. We train a public POINTS-1.5 model to obtain POINTS-Reader, which surpasses many existing public and proprietary models of comparable or larger size. Our model is available at https://github.com/Tencent/POINTS-Reader.
title POINTS-Reader: Distillation-Free Adaptation of Vision-Language Models for Document Conversion
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.01215