PaddleOCR-VL-1.5: Towards a Multi-Task 0.9B VLM for Robust In-the-Wild Document Parsing

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cui, Cheng, Sun, Ting, Liang, Suyin, Gao, Tingquan, Zhang, Zelun, Liu, Jiaxuan, Wang, Xueqing, Zhou, Changda, Liu, Hongen, Lin, Manhui, Zhang, Yue, Zhang, Yubo, Liu, Yi, Yu, Dianhai, Ma, Yanjun
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918426002325504
author Cui, Cheng
Sun, Ting
Liang, Suyin
Gao, Tingquan
Zhang, Zelun
Liu, Jiaxuan
Wang, Xueqing
Zhou, Changda
Liu, Hongen
Lin, Manhui
Zhang, Yue
Zhang, Yubo
Liu, Yi
Yu, Dianhai
Ma, Yanjun
author_facet Cui, Cheng
Sun, Ting
Liang, Suyin
Gao, Tingquan
Zhang, Zelun
Liu, Jiaxuan
Wang, Xueqing
Zhou, Changda
Liu, Hongen
Lin, Manhui
Zhang, Yue
Zhang, Yubo
Liu, Yi
Yu, Dianhai
Ma, Yanjun
contents We introduce PaddleOCR-VL-1.5, an upgraded model achieving a new state-of-the-art (SOTA) accuracy of 94.5% on OmniDocBench v1.5. To rigorously evaluate robustness against real-world physical distortions, including scanning, skew, warping, screen-photography, and illumination, we propose the Real5-OmniDocBench benchmark. Experimental results demonstrate that this enhanced model attains SOTA performance on the newly curated benchmark. Furthermore, we extend the model's capabilities by incorporating seal recognition and text spotting tasks, while remaining a 0.9B ultra-compact VLM with high efficiency. Code: https://github.com/PaddlePaddle/PaddleOCR
format Preprint
id arxiv_https___arxiv_org_abs_2601_21957
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle PaddleOCR-VL-1.5: Towards a Multi-Task 0.9B VLM for Robust In-the-Wild Document Parsing
Cui, Cheng
Sun, Ting
Liang, Suyin
Gao, Tingquan
Zhang, Zelun
Liu, Jiaxuan
Wang, Xueqing
Zhou, Changda
Liu, Hongen
Lin, Manhui
Zhang, Yue
Zhang, Yubo
Liu, Yi
Yu, Dianhai
Ma, Yanjun
Computer Vision and Pattern Recognition
We introduce PaddleOCR-VL-1.5, an upgraded model achieving a new state-of-the-art (SOTA) accuracy of 94.5% on OmniDocBench v1.5. To rigorously evaluate robustness against real-world physical distortions, including scanning, skew, warping, screen-photography, and illumination, we propose the Real5-OmniDocBench benchmark. Experimental results demonstrate that this enhanced model attains SOTA performance on the newly curated benchmark. Furthermore, we extend the model's capabilities by incorporating seal recognition and text spotting tasks, while remaining a 0.9B ultra-compact VLM with high efficiency. Code: https://github.com/PaddlePaddle/PaddleOCR
title PaddleOCR-VL-1.5: Towards a Multi-Task 0.9B VLM for Robust In-the-Wild Document Parsing
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2601.21957