PaddleOCR 3.0 Technical Report
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Cui, Cheng, Sun, Ting, Lin, Manhui, Gao, Tingquan, Zhang, Yubo, Liu, Jiaxuan, Wang, Xueqing, Zhang, Zelun, Zhou, Changda, Liu, Hongen, Zhang, Yue, Lv, Wenyu, Huang, Kui, Zhang, Yichao, Zhang, Jing, Zhang, Jun, Liu, Yi, Yu, Dianhai, Ma, Yanjun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
PaddleOCR-VL-1.5: Towards a Multi-Task 0.9B VLM for Robust In-the-Wild Document Parsing
von: Cui, Cheng, et al.
Veröffentlicht: (2026)
von: Cui, Cheng, et al.
Veröffentlicht: (2026)
PaddleOCR-VL: Boosting Multilingual Document Parsing via a 0.9B Ultra-Compact Vision-Language Model
von: Cui, Cheng, et al.
Veröffentlicht: (2025)
von: Cui, Cheng, et al.
Veröffentlicht: (2025)
PP-OCRv5: A Specialized 5M-Parameter Model Rivaling Billion-Parameter Vision-Language Models on OCR Tasks
von: Cui, Cheng, et al.
Veröffentlicht: (2026)
von: Cui, Cheng, et al.
Veröffentlicht: (2026)
Boosting Document Parsing Efficiency and Performance with Coarse-to-Fine Visual Processing
von: Cui, Cheng, et al.
Veröffentlicht: (2026)
von: Cui, Cheng, et al.
Veröffentlicht: (2026)
A Novel Implementation of Marksheet Parser Using PaddleOCR
von: Bagaria, Sankalp, et al.
Veröffentlicht: (2024)
von: Bagaria, Sankalp, et al.
Veröffentlicht: (2024)
Real5-OmniDocBench: A Full-Scale Physical Reconstruction Benchmark for Robust Document Parsing in the Wild
von: Zhou, Changda, et al.
Veröffentlicht: (2026)
von: Zhou, Changda, et al.
Veröffentlicht: (2026)
GLM-OCR Technical Report
von: Duan, Shuaiqi, et al.
Veröffentlicht: (2026)
von: Duan, Shuaiqi, et al.
Veröffentlicht: (2026)
Augmenting Human Cognition With Generative AI: Lessons From AI-Assisted Decision-Making
von: Zhang, Zelun Tony, et al.
Veröffentlicht: (2025)
von: Zhang, Zelun Tony, et al.
Veröffentlicht: (2025)
HunyuanImage 3.0 Technical Report
von: Cao, Siyu, et al.
Veröffentlicht: (2025)
von: Cao, Siyu, et al.
Veröffentlicht: (2025)
HunyuanOCR Technical Report
von: Hunyuan Vision Team, et al.
Veröffentlicht: (2025)
von: Hunyuan Vision Team, et al.
Veröffentlicht: (2025)
Carbon and Reliability-Aware Computing for Heterogeneous Data Centers
von: Zhang, Yichao, et al.
Veröffentlicht: (2025)
von: Zhang, Yichao, et al.
Veröffentlicht: (2025)
A Monkey‐Saddle‐Shaped Nanographene Embedding a Dicyclohepta[ cd , fg ]‐ as ‐Indacene Core and Two Additional Heptagons, and Its Paddle‐Wheel Co‐Assembly With Fullerenes
von: Wenjun Liu, et al.
Veröffentlicht: (2026)
von: Wenjun Liu, et al.
Veröffentlicht: (2026)
MonkeyOCR v1.5 Technical Report: Unlocking Robust Document Parsing for Complex Patterns
von: Zhang, Jiarui, et al.
Veröffentlicht: (2025)
von: Zhang, Jiarui, et al.
Veröffentlicht: (2025)
OmniOCR: Generalist OCR for Ethnic Minority Languages
von: Liu, Bonan, et al.
Veröffentlicht: (2026)
von: Liu, Bonan, et al.
Veröffentlicht: (2026)
Seedream 3.0 Technical Report
von: Gao, Yu, et al.
Veröffentlicht: (2025)
von: Gao, Yu, et al.
Veröffentlicht: (2025)
OCR-Quality: A Human-Annotated Dataset for OCR Quality Assessment
von: Zhang, Yulong
Veröffentlicht: (2025)
von: Zhang, Yulong
Veröffentlicht: (2025)
Task Allocation Algorithm and Simulation Analysis for Multiple AMRs in Digital‐Intelligent Warehouses
von: Zixia Chen, et al.
Veröffentlicht: (2025)
von: Zixia Chen, et al.
Veröffentlicht: (2025)
Heterometallic Nickel(II) Diruthenium(II, III) Carbonates by Self‐Assembly of Paddle‐Wheel Precursors: Insights into Structures and Magnetic Properties
von: Qian‐Qian Liang, et al.
Veröffentlicht: (2026)
von: Qian‐Qian Liang, et al.
Veröffentlicht: (2026)
Magnetic nanofluid‐based liquid marble for a self‐powered mechanosensation
von: Manhui Chen, et al.
Veröffentlicht: (2024)
von: Manhui Chen, et al.
Veröffentlicht: (2024)
TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document
von: Liu, Yuliang, et al.
Veröffentlicht: (2024)
von: Liu, Yuliang, et al.
Veröffentlicht: (2024)
MonkeyOCR: Document Parsing with a Structure-Recognition-Relation Triplet Paradigm
von: Li, Zhang, et al.
Veröffentlicht: (2025)
von: Li, Zhang, et al.
Veröffentlicht: (2025)
Shot Segmentation Based on Von Neumann Entropy for Key Frame Extraction
von: Zhang, Xueqing, et al.
Veröffentlicht: (2024)
von: Zhang, Xueqing, et al.
Veröffentlicht: (2024)
Summing to Uncertainty: On the Necessity of Additivity in Deriving the Born Rule
von: Zhang, Jiaxuan
Veröffentlicht: (2026)
von: Zhang, Jiaxuan
Veröffentlicht: (2026)
Sleeping Beauty in One or Many Worlds: A Defense of the Halfer Position
von: Zhang, Jiaxuan
Veröffentlicht: (2026)
von: Zhang, Jiaxuan
Veröffentlicht: (2026)
Enhancing OCR for Sino-Vietnamese Language Processing via Fine-tuned PaddleOCRv5
von: Nguyen, Minh Hoang, et al.
Veröffentlicht: (2025)
von: Nguyen, Minh Hoang, et al.
Veröffentlicht: (2025)
MME-VideoOCR: Evaluating OCR-Based Capabilities of Multimodal LLMs in Video Scenarios
von: Shi, Yang, et al.
Veröffentlicht: (2025)
von: Shi, Yang, et al.
Veröffentlicht: (2025)
Hilbert Poincaré series and kernels for products of $L$-functions
von: Zhang, Mingkuan, et al.
Veröffentlicht: (2024)
von: Zhang, Mingkuan, et al.
Veröffentlicht: (2024)
Racing Against a Career‐Fertility Countdown: The Prospective Motherhood Penalty and Gendered Ageism in China's Workplace
von: Rose Xueqing Zhang
Veröffentlicht: (2025)
von: Rose Xueqing Zhang
Veröffentlicht: (2025)
Constructing Built‐In Electric Field in Na 3 V 2 (PO 4 ) 3 /NaV(P 2 O 7 ) Heterostructure With Columnar Cluster Morphology Boosting High Capacity and Energy Density for Sodium Ion Batteries
von: Shuming Zhang, et al.
Veröffentlicht: (2025)
von: Shuming Zhang, et al.
Veröffentlicht: (2025)
Lung cancer dose distribution prediction based on a dual‐branch feature extraction network
von: Haifeng Zhang, et al.
Veröffentlicht: (2025)
von: Haifeng Zhang, et al.
Veröffentlicht: (2025)
Ultra-broadband acoustic absorber based on periodic acoustic rigid-metaporous composite array
von: Zhang, Dongguo, et al.
Veröffentlicht: (2024)
von: Zhang, Dongguo, et al.
Veröffentlicht: (2024)
OCR Hinders RAG: Evaluating the Cascading Impact of OCR on Retrieval-Augmented Generation
von: Zhang, Junyuan, et al.
Veröffentlicht: (2024)
von: Zhang, Junyuan, et al.
Veröffentlicht: (2024)
Structural dynamics of plant transcription factors and their functional implications
von: Shaowen Wu, et al.
Veröffentlicht: (2026)
von: Shaowen Wu, et al.
Veröffentlicht: (2026)
An anti‐windup and backlash compensation‐based finite‐time control method for performance enhancement of a class of nonlinear systems
von: Guangyu Liu, et al.
Veröffentlicht: (2024)
von: Guangyu Liu, et al.
Veröffentlicht: (2024)
FireRed-OCR Technical Report
von: Wu, Hao, et al.
Veröffentlicht: (2026)
von: Wu, Hao, et al.
Veröffentlicht: (2026)
OCRGenBench: A Comprehensive Benchmark for Evaluating OCR Generative Capabilities
von: Zhang, Peirong, et al.
Veröffentlicht: (2025)
von: Zhang, Peirong, et al.
Veröffentlicht: (2025)
MemOCR: Layout-Aware Visual Memory for Efficient Long-Horizon Reasoning
von: Shi, Yaorui, et al.
Veröffentlicht: (2026)
von: Shi, Yaorui, et al.
Veröffentlicht: (2026)
Agentar-Fin-OCR
von: Qian, Siyi, et al.
Veröffentlicht: (2026)
von: Qian, Siyi, et al.
Veröffentlicht: (2026)
General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model
von: Wei, Haoran, et al.
Veröffentlicht: (2024)
von: Wei, Haoran, et al.
Veröffentlicht: (2024)
When Good OCR Is Not Enough: Benchmarking OCR Robustness for Retrieval-Augmented Generation
von: Sun, Lin, et al.
Veröffentlicht: (2026)
von: Sun, Lin, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
PaddleOCR-VL-1.5: Towards a Multi-Task 0.9B VLM for Robust In-the-Wild Document Parsing
von: Cui, Cheng, et al.
Veröffentlicht: (2026) -
PaddleOCR-VL: Boosting Multilingual Document Parsing via a 0.9B Ultra-Compact Vision-Language Model
von: Cui, Cheng, et al.
Veröffentlicht: (2025) -
PP-OCRv5: A Specialized 5M-Parameter Model Rivaling Billion-Parameter Vision-Language Models on OCR Tasks
von: Cui, Cheng, et al.
Veröffentlicht: (2026) -
Boosting Document Parsing Efficiency and Performance with Coarse-to-Fine Visual Processing
von: Cui, Cheng, et al.
Veröffentlicht: (2026) -
A Novel Implementation of Marksheet Parser Using PaddleOCR
von: Bagaria, Sankalp, et al.
Veröffentlicht: (2024)