Qianfan-OCR: A Unified End-to-End Model for Document Intelligence
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911514591494144 |
|---|---|
| author | Dong, Daxiang Zheng, Mingming Xu, Dong Luo, Chunhua Zhuang, Bairong Li, Yuxuan He, Ruoyun Wang, Haoran Zhang, Wenyu Wang, Wenbo Wang, Yicheng Xiong, Xue Zheng, Ayong Zuo, Xiaoying Ou, Ziwei Gu, Jingnan Guo, Quanhao Wu, Jianmin Yin, Dawei Shen, Dou |
| author_facet | Dong, Daxiang Zheng, Mingming Xu, Dong Luo, Chunhua Zhuang, Bairong Li, Yuxuan He, Ruoyun Wang, Haoran Zhang, Wenyu Wang, Wenbo Wang, Yicheng Xiong, Xue Zheng, Ayong Zuo, Xiaoying Ou, Ziwei Gu, Jingnan Guo, Quanhao Wu, Jianmin Yin, Dawei Shen, Dou |
| contents | We present Qianfan-OCR, a 4B-parameter end-to-end vision-language model that unifies document parsing, layout analysis, and document understanding within a single architecture. It performs direct image-to-Markdown conversion and supports diverse prompt-driven tasks including table extraction, chart understanding, document QA, and key information extraction. To address the loss of explicit layout analysis in end-to-end OCR, we propose Layout-as-Thought, an optional thinking phase triggered by special think tokens that generates structured layout representations -- bounding boxes, element types, and reading order -- before producing final outputs, recovering layout grounding capabilities while improving accuracy on complex layouts. Qianfan-OCR ranks first among end-to-end models on OmniDocBench v1.5 (93.12) and OlmOCR Bench (79.8), achieves competitive results on OCRBench, CCOCR, DocVQA, and ChartQA against general VLMs of comparable scale, and attains the highest average score on public key information extraction benchmarks, surpassing Gemini-3.1-Pro, Seed-2.0, and Qwen3-VL-235B. The model is publicly accessible via the Baidu AI Cloud Qianfan platform. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2603_13398 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Qianfan-OCR: A Unified End-to-End Model for Document Intelligence Dong, Daxiang Zheng, Mingming Xu, Dong Luo, Chunhua Zhuang, Bairong Li, Yuxuan He, Ruoyun Wang, Haoran Zhang, Wenyu Wang, Wenbo Wang, Yicheng Xiong, Xue Zheng, Ayong Zuo, Xiaoying Ou, Ziwei Gu, Jingnan Guo, Quanhao Wu, Jianmin Yin, Dawei Shen, Dou Computer Vision and Pattern Recognition We present Qianfan-OCR, a 4B-parameter end-to-end vision-language model that unifies document parsing, layout analysis, and document understanding within a single architecture. It performs direct image-to-Markdown conversion and supports diverse prompt-driven tasks including table extraction, chart understanding, document QA, and key information extraction. To address the loss of explicit layout analysis in end-to-end OCR, we propose Layout-as-Thought, an optional thinking phase triggered by special think tokens that generates structured layout representations -- bounding boxes, element types, and reading order -- before producing final outputs, recovering layout grounding capabilities while improving accuracy on complex layouts. Qianfan-OCR ranks first among end-to-end models on OmniDocBench v1.5 (93.12) and OlmOCR Bench (79.8), achieves competitive results on OCRBench, CCOCR, DocVQA, and ChartQA against general VLMs of comparable scale, and attains the highest average score on public key information extraction benchmarks, surpassing Gemini-3.1-Pro, Seed-2.0, and Qwen3-VL-235B. The model is publicly accessible via the Baidu AI Cloud Qianfan platform. |
| title | Qianfan-OCR: A Unified End-to-End Model for Document Intelligence |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2603.13398 |