PaddleOCR-VL-1.5: Towards a Multi-Task 0.9B VLM for Robust In-the-Wild Document Parsing
Fuente:
arXiv
Saved in:
| Main Authors: | Cui, Cheng, Sun, Ting, Liang, Suyin, Gao, Tingquan, Zhang, Zelun, Liu, Jiaxuan, Wang, Xueqing, Zhou, Changda, Liu, Hongen, Lin, Manhui, Zhang, Yue, Zhang, Yubo, Liu, Yi, Yu, Dianhai, Ma, Yanjun |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PaddleOCR-VL: Boosting Multilingual Document Parsing via a 0.9B Ultra-Compact Vision-Language Model
by: Cui, Cheng, et al.
Published: (2025)
by: Cui, Cheng, et al.
Published: (2025)
PaddleOCR 3.0 Technical Report
by: Cui, Cheng, et al.
Published: (2025)
by: Cui, Cheng, et al.
Published: (2025)
Boosting Document Parsing Efficiency and Performance with Coarse-to-Fine Visual Processing
by: Cui, Cheng, et al.
Published: (2026)
by: Cui, Cheng, et al.
Published: (2026)
PP-OCRv5: A Specialized 5M-Parameter Model Rivaling Billion-Parameter Vision-Language Models on OCR Tasks
by: Cui, Cheng, et al.
Published: (2026)
by: Cui, Cheng, et al.
Published: (2026)
Real5-OmniDocBench: A Full-Scale Physical Reconstruction Benchmark for Robust Document Parsing in the Wild
by: Zhou, Changda, et al.
Published: (2026)
by: Zhou, Changda, et al.
Published: (2026)
A Novel Implementation of Marksheet Parser Using PaddleOCR
by: Bagaria, Sankalp, et al.
Published: (2024)
by: Bagaria, Sankalp, et al.
Published: (2024)
Multimodal OCR: Parse Anything from Documents
by: Zheng, Handong, et al.
Published: (2026)
by: Zheng, Handong, et al.
Published: (2026)
MonkeyOCR: Document Parsing with a Structure-Recognition-Relation Triplet Paradigm
by: Li, Zhang, et al.
Published: (2025)
by: Li, Zhang, et al.
Published: (2025)
MonkeyOCR v1.5 Technical Report: Unlocking Robust Document Parsing for Complex Patterns
by: Zhang, Jiarui, et al.
Published: (2025)
by: Zhang, Jiarui, et al.
Published: (2025)
DOCR-Inspector: Fine-Grained and Automated Evaluation of Document Parsing with VLM
by: Zhang, Qintong, et al.
Published: (2025)
by: Zhang, Qintong, et al.
Published: (2025)
FastOCR: Dynamic Visual Fixation via KV Cache Pruning for Efficient Document Parsing
by: Tang, Zihan, et al.
Published: (2026)
by: Tang, Zihan, et al.
Published: (2026)
Consensus Entropy: Harnessing Multi-VLM Agreement for Self-Verifying and Self-Improving OCR
by: Zhang, Yulong, et al.
Published: (2025)
by: Zhang, Yulong, et al.
Published: (2025)
MeDocVL: A Visual Language Model for Medical Document Understanding and Parsing
by: Wang, Wenjie, et al.
Published: (2026)
by: Wang, Wenjie, et al.
Published: (2026)
Task Allocation Algorithm and Simulation Analysis for Multiple AMRs in Digital‐Intelligent Warehouses
by: Zixia Chen, et al.
Published: (2025)
by: Zixia Chen, et al.
Published: (2025)
DocFusion: A Unified Framework for Document Parsing Tasks
by: Chai, Mingxu, et al.
Published: (2024)
by: Chai, Mingxu, et al.
Published: (2024)
MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding
by: Xie, Rongchang, et al.
Published: (2024)
by: Xie, Rongchang, et al.
Published: (2024)
TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document
by: Liu, Yuliang, et al.
Published: (2024)
by: Liu, Yuliang, et al.
Published: (2024)
MDPBench: A Benchmark for Multilingual Document Parsing in Real-World Scenarios
by: Li, Zhang, et al.
Published: (2026)
by: Li, Zhang, et al.
Published: (2026)
Enhancing OCR for Sino-Vietnamese Language Processing via Fine-tuned PaddleOCRv5
by: Nguyen, Minh Hoang, et al.
Published: (2025)
by: Nguyen, Minh Hoang, et al.
Published: (2025)
Augmenting Human Cognition With Generative AI: Lessons From AI-Assisted Decision-Making
by: Zhang, Zelun Tony, et al.
Published: (2025)
by: Zhang, Zelun Tony, et al.
Published: (2025)
Penguin-VL: Exploring the Efficiency Limits of VLM with LLM-based Vision Encoders
by: Zhang, Boqiang, et al.
Published: (2026)
by: Zhang, Boqiang, et al.
Published: (2026)
Seed1.5-VL Technical Report
by: Guo, Dong, et al.
Published: (2025)
by: Guo, Dong, et al.
Published: (2025)
Letter to the Editor: Effect of Aerobic Combined Resistance Exercise in Dialysis on Restless Legs Syndrome: A Randomized Controlled Study
by: Suyin Yu
Published: (2026)
by: Suyin Yu
Published: (2026)
The morning deluge : Mao Tsetung and the chinese revolution, 1893-1954 / Han Suyin
by: Suyin, Han
Published: (1972)
by: Suyin, Han
Published: (1972)
A Monkey‐Saddle‐Shaped Nanographene Embedding a Dicyclohepta[ cd , fg ]‐ as ‐Indacene Core and Two Additional Heptagons, and Its Paddle‐Wheel Co‐Assembly With Fullerenes
by: Wenjun Liu, et al.
Published: (2026)
by: Wenjun Liu, et al.
Published: (2026)
OmniOCR: Generalist OCR for Ethnic Minority Languages
by: Liu, Bonan, et al.
Published: (2026)
by: Liu, Bonan, et al.
Published: (2026)
DocParseNet: Advanced Semantic Segmentation and OCR Embeddings for Efficient Scanned Document Annotation
by: Mohammadshirazi, Ahmad, et al.
Published: (2024)
by: Mohammadshirazi, Ahmad, et al.
Published: (2024)
Singpath-VL Technical Report
by: Qiu, Zhen, et al.
Published: (2026)
by: Qiu, Zhen, et al.
Published: (2026)
AgenticOCR: Parsing Only What You Need for Efficient Retrieval-Augmented Generation
by: Wang, Zhengren, et al.
Published: (2026)
by: Wang, Zhengren, et al.
Published: (2026)
WebAccessVL: Violation-Aware VLM for Web Accessibility
by: Zheng, Amber Yijia, et al.
Published: (2025)
by: Zheng, Amber Yijia, et al.
Published: (2025)
PP-FormulaNet: Bridging Accuracy and Efficiency in Advanced Formula Recognition
by: Liu, Hongen, et al.
Published: (2025)
by: Liu, Hongen, et al.
Published: (2025)
dots.ocr: Multilingual Document Layout Parsing in a Single Vision-Language Model
by: Li, Yumeng, et al.
Published: (2025)
by: Li, Yumeng, et al.
Published: (2025)
ParseBench: A Document Parsing Benchmark for AI Agents
by: Zhang, Boyang, et al.
Published: (2026)
by: Zhang, Boyang, et al.
Published: (2026)
mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding
by: Hu, Anwen, et al.
Published: (2024)
by: Hu, Anwen, et al.
Published: (2024)
Parser-Oriented Structural Refinement for a Stable Layout Interface in Document Parsing
by: Liu, Fuyuan, et al.
Published: (2026)
by: Liu, Fuyuan, et al.
Published: (2026)
Kwai Keye-VL 1.5 Technical Report
by: Yang, Biao, et al.
Published: (2025)
by: Yang, Biao, et al.
Published: (2025)
Heterometallic Nickel(II) Diruthenium(II, III) Carbonates by Self‐Assembly of Paddle‐Wheel Precursors: Insights into Structures and Magnetic Properties
by: Qian‐Qian Liang, et al.
Published: (2026)
by: Qian‐Qian Liang, et al.
Published: (2026)
MGFFD-VLM: Multi-Granularity Prompt Learning for Face Forgery Detection with VLM
by: Chen, Tao, et al.
Published: (2025)
by: Chen, Tao, et al.
Published: (2025)
Radar Target Detection with K‐Nearest Neighbor Manifold Filter on Riemannian Manifold
by: Dongao Zhou, et al.
Published: (2024)
by: Dongao Zhou, et al.
Published: (2024)
Large-Scale Constraint Generation -- Can LLMs Parse Hundreds of Constraints?
by: Boffa, Matteo, et al.
Published: (2025)
by: Boffa, Matteo, et al.
Published: (2025)
Similar Items
-
PaddleOCR-VL: Boosting Multilingual Document Parsing via a 0.9B Ultra-Compact Vision-Language Model
by: Cui, Cheng, et al.
Published: (2025) -
PaddleOCR 3.0 Technical Report
by: Cui, Cheng, et al.
Published: (2025) -
Boosting Document Parsing Efficiency and Performance with Coarse-to-Fine Visual Processing
by: Cui, Cheng, et al.
Published: (2026) -
PP-OCRv5: A Specialized 5M-Parameter Model Rivaling Billion-Parameter Vision-Language Models on OCR Tasks
by: Cui, Cheng, et al.
Published: (2026) -
Real5-OmniDocBench: A Full-Scale Physical Reconstruction Benchmark for Robust Document Parsing in the Wild
by: Zhou, Changda, et al.
Published: (2026)