PaddleOCR-VL: Boosting Multilingual Document Parsing via a 0.9B Ultra-Compact Vision-Language Model
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Cui, Cheng, Sun, Ting, Liang, Suyin, Gao, Tingquan, Zhang, Zelun, Liu, Jiaxuan, Wang, Xueqing, Zhou, Changda, Liu, Hongen, Lin, Manhui, Zhang, Yue, Zhang, Yubo, Zheng, Handong, Zhang, Jing, Zhang, Jun, Liu, Yi, Yu, Dianhai, Ma, Yanjun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
PaddleOCR-VL-1.5: Towards a Multi-Task 0.9B VLM for Robust In-the-Wild Document Parsing
von: Cui, Cheng, et al.
Veröffentlicht: (2026)
von: Cui, Cheng, et al.
Veröffentlicht: (2026)
PaddleOCR 3.0 Technical Report
von: Cui, Cheng, et al.
Veröffentlicht: (2025)
von: Cui, Cheng, et al.
Veröffentlicht: (2025)
Boosting Document Parsing Efficiency and Performance with Coarse-to-Fine Visual Processing
von: Cui, Cheng, et al.
Veröffentlicht: (2026)
von: Cui, Cheng, et al.
Veröffentlicht: (2026)
PP-OCRv5: A Specialized 5M-Parameter Model Rivaling Billion-Parameter Vision-Language Models on OCR Tasks
von: Cui, Cheng, et al.
Veröffentlicht: (2026)
von: Cui, Cheng, et al.
Veröffentlicht: (2026)
A Novel Implementation of Marksheet Parser Using PaddleOCR
von: Bagaria, Sankalp, et al.
Veröffentlicht: (2024)
von: Bagaria, Sankalp, et al.
Veröffentlicht: (2024)
Real5-OmniDocBench: A Full-Scale Physical Reconstruction Benchmark for Robust Document Parsing in the Wild
von: Zhou, Changda, et al.
Veröffentlicht: (2026)
von: Zhou, Changda, et al.
Veröffentlicht: (2026)
Multimodal OCR: Parse Anything from Documents
von: Zheng, Handong, et al.
Veröffentlicht: (2026)
von: Zheng, Handong, et al.
Veröffentlicht: (2026)
MonkeyOCR: Document Parsing with a Structure-Recognition-Relation Triplet Paradigm
von: Li, Zhang, et al.
Veröffentlicht: (2025)
von: Li, Zhang, et al.
Veröffentlicht: (2025)
MDPBench: A Benchmark for Multilingual Document Parsing in Real-World Scenarios
von: Li, Zhang, et al.
Veröffentlicht: (2026)
von: Li, Zhang, et al.
Veröffentlicht: (2026)
Compact SPICE model for TeraFET resonant detectors
von: Liu, Xueqing, et al.
Veröffentlicht: (2024)
von: Liu, Xueqing, et al.
Veröffentlicht: (2024)
MonkeyOCR v1.5 Technical Report: Unlocking Robust Document Parsing for Complex Patterns
von: Zhang, Jiarui, et al.
Veröffentlicht: (2025)
von: Zhang, Jiarui, et al.
Veröffentlicht: (2025)
Augmenting Human Cognition With Generative AI: Lessons From AI-Assisted Decision-Making
von: Zhang, Zelun Tony, et al.
Veröffentlicht: (2025)
von: Zhang, Zelun Tony, et al.
Veröffentlicht: (2025)
dots.ocr: Multilingual Document Layout Parsing in a Single Vision-Language Model
von: Li, Yumeng, et al.
Veröffentlicht: (2025)
von: Li, Yumeng, et al.
Veröffentlicht: (2025)
Constructing Built‐In Electric Field in Na 3 V 2 (PO 4 ) 3 /NaV(P 2 O 7 ) Heterostructure With Columnar Cluster Morphology Boosting High Capacity and Energy Density for Sodium Ion Batteries
von: Shuming Zhang, et al.
Veröffentlicht: (2025)
von: Shuming Zhang, et al.
Veröffentlicht: (2025)
A Monkey‐Saddle‐Shaped Nanographene Embedding a Dicyclohepta[ cd , fg ]‐ as ‐Indacene Core and Two Additional Heptagons, and Its Paddle‐Wheel Co‐Assembly With Fullerenes
von: Wenjun Liu, et al.
Veröffentlicht: (2026)
von: Wenjun Liu, et al.
Veröffentlicht: (2026)
FastOCR: Dynamic Visual Fixation via KV Cache Pruning for Efficient Document Parsing
von: Tang, Zihan, et al.
Veröffentlicht: (2026)
von: Tang, Zihan, et al.
Veröffentlicht: (2026)
OmniOCR: Generalist OCR for Ethnic Minority Languages
von: Liu, Bonan, et al.
Veröffentlicht: (2026)
von: Liu, Bonan, et al.
Veröffentlicht: (2026)
Task Allocation Algorithm and Simulation Analysis for Multiple AMRs in Digital‐Intelligent Warehouses
von: Zixia Chen, et al.
Veröffentlicht: (2025)
von: Zixia Chen, et al.
Veröffentlicht: (2025)
Singpath-VL Technical Report
von: Qiu, Zhen, et al.
Veröffentlicht: (2026)
von: Qiu, Zhen, et al.
Veröffentlicht: (2026)
Explore the Limits of Omni-modal Pretraining at Scale
von: Zhang, Yiyuan, et al.
Veröffentlicht: (2024)
von: Zhang, Yiyuan, et al.
Veröffentlicht: (2024)
AgenticOCR: Parsing Only What You Need for Efficient Retrieval-Augmented Generation
von: Wang, Zhengren, et al.
Veröffentlicht: (2026)
von: Wang, Zhengren, et al.
Veröffentlicht: (2026)
Enhancing OCR for Sino-Vietnamese Language Processing via Fine-tuned PaddleOCRv5
von: Nguyen, Minh Hoang, et al.
Veröffentlicht: (2025)
von: Nguyen, Minh Hoang, et al.
Veröffentlicht: (2025)
OCR-Quality: A Human-Annotated Dataset for OCR Quality Assessment
von: Zhang, Yulong
Veröffentlicht: (2025)
von: Zhang, Yulong
Veröffentlicht: (2025)
Heterometallic Nickel(II) Diruthenium(II, III) Carbonates by Self‐Assembly of Paddle‐Wheel Precursors: Insights into Structures and Magnetic Properties
von: Qian‐Qian Liang, et al.
Veröffentlicht: (2026)
von: Qian‐Qian Liang, et al.
Veröffentlicht: (2026)
Magnetic nanofluid‐based liquid marble for a self‐powered mechanosensation
von: Manhui Chen, et al.
Veröffentlicht: (2024)
von: Manhui Chen, et al.
Veröffentlicht: (2024)
TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document
von: Liu, Yuliang, et al.
Veröffentlicht: (2024)
von: Liu, Yuliang, et al.
Veröffentlicht: (2024)
Shot Segmentation Based on Von Neumann Entropy for Key Frame Extraction
von: Zhang, Xueqing, et al.
Veröffentlicht: (2024)
von: Zhang, Xueqing, et al.
Veröffentlicht: (2024)
MeDocVL: A Visual Language Model for Medical Document Understanding and Parsing
von: Wang, Wenjie, et al.
Veröffentlicht: (2026)
von: Wang, Wenjie, et al.
Veröffentlicht: (2026)
Graph of Records: Boosting Retrieval Augmented Generation for Long-context Summarization with Graphs
von: Zhang, Haozhen, et al.
Veröffentlicht: (2024)
von: Zhang, Haozhen, et al.
Veröffentlicht: (2024)
Letter to the Editor: Effect of Aerobic Combined Resistance Exercise in Dialysis on Restless Legs Syndrome: A Randomized Controlled Study
von: Suyin Yu
Veröffentlicht: (2026)
von: Suyin Yu
Veröffentlicht: (2026)
The morning deluge : Mao Tsetung and the chinese revolution, 1893-1954 / Han Suyin
von: Suyin, Han
Veröffentlicht: (1972)
von: Suyin, Han
Veröffentlicht: (1972)
Unique Oxygen‐Bridged Nickel Atomic Pairs Efficiently Boost Electrochemical Reduction of Carbon Dioxide
von: Chaofan Zhang, et al.
Veröffentlicht: (2024)
von: Chaofan Zhang, et al.
Veröffentlicht: (2024)
MME-VideoOCR: Evaluating OCR-Based Capabilities of Multimodal LLMs in Video Scenarios
von: Shi, Yang, et al.
Veröffentlicht: (2025)
von: Shi, Yang, et al.
Veröffentlicht: (2025)
Multilingual Multi-Aspect Explainability Analyses on Machine Reading Comprehension Models
von: Cui, Yiming, et al.
Veröffentlicht: (2021)
von: Cui, Yiming, et al.
Veröffentlicht: (2021)
Vid-LLM: A Compact Video-based 3D Multimodal LLM with Reconstruction-Reasoning Synergy
von: Chen, Haijier, et al.
Veröffentlicht: (2025)
von: Chen, Haijier, et al.
Veröffentlicht: (2025)
Summing to Uncertainty: On the Necessity of Additivity in Deriving the Born Rule
von: Zhang, Jiaxuan
Veröffentlicht: (2026)
von: Zhang, Jiaxuan
Veröffentlicht: (2026)
Sleeping Beauty in One or Many Worlds: A Defense of the Halfer Position
von: Zhang, Jiaxuan
Veröffentlicht: (2026)
von: Zhang, Jiaxuan
Veröffentlicht: (2026)
SciFlow-Bench: Evaluating Structure-Aware Scientific Diagram Generation via Inverse Parsing
von: Zhang, Tong, et al.
Veröffentlicht: (2026)
von: Zhang, Tong, et al.
Veröffentlicht: (2026)
Racing Against a Career‐Fertility Countdown: The Prospective Motherhood Penalty and Gendered Ageism in China's Workplace
von: Rose Xueqing Zhang
Veröffentlicht: (2025)
von: Rose Xueqing Zhang
Veröffentlicht: (2025)
Structural dynamics of plant transcription factors and their functional implications
von: Shaowen Wu, et al.
Veröffentlicht: (2026)
von: Shaowen Wu, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
PaddleOCR-VL-1.5: Towards a Multi-Task 0.9B VLM for Robust In-the-Wild Document Parsing
von: Cui, Cheng, et al.
Veröffentlicht: (2026) -
PaddleOCR 3.0 Technical Report
von: Cui, Cheng, et al.
Veröffentlicht: (2025) -
Boosting Document Parsing Efficiency and Performance with Coarse-to-Fine Visual Processing
von: Cui, Cheng, et al.
Veröffentlicht: (2026) -
PP-OCRv5: A Specialized 5M-Parameter Model Rivaling Billion-Parameter Vision-Language Models on OCR Tasks
von: Cui, Cheng, et al.
Veröffentlicht: (2026) -
A Novel Implementation of Marksheet Parser Using PaddleOCR
von: Bagaria, Sankalp, et al.
Veröffentlicht: (2024)