Qianfan-OCR: A Unified End-to-End Model for Document Intelligence

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dong, Daxiang, Zheng, Mingming, Xu, Dong, Luo, Chunhua, Zhuang, Bairong, Li, Yuxuan, He, Ruoyun, Wang, Haoran, Zhang, Wenyu, Wang, Wenbo, Wang, Yicheng, Xiong, Xue, Zheng, Ayong, Zuo, Xiaoying, Ou, Ziwei, Gu, Jingnan, Guo, Quanhao, Wu, Jianmin, Yin, Dawei, Shen, Dou
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911514591494144
author Dong, Daxiang
Zheng, Mingming
Xu, Dong
Luo, Chunhua
Zhuang, Bairong
Li, Yuxuan
He, Ruoyun
Wang, Haoran
Zhang, Wenyu
Wang, Wenbo
Wang, Yicheng
Xiong, Xue
Zheng, Ayong
Zuo, Xiaoying
Ou, Ziwei
Gu, Jingnan
Guo, Quanhao
Wu, Jianmin
Yin, Dawei
Shen, Dou
author_facet Dong, Daxiang
Zheng, Mingming
Xu, Dong
Luo, Chunhua
Zhuang, Bairong
Li, Yuxuan
He, Ruoyun
Wang, Haoran
Zhang, Wenyu
Wang, Wenbo
Wang, Yicheng
Xiong, Xue
Zheng, Ayong
Zuo, Xiaoying
Ou, Ziwei
Gu, Jingnan
Guo, Quanhao
Wu, Jianmin
Yin, Dawei
Shen, Dou
contents We present Qianfan-OCR, a 4B-parameter end-to-end vision-language model that unifies document parsing, layout analysis, and document understanding within a single architecture. It performs direct image-to-Markdown conversion and supports diverse prompt-driven tasks including table extraction, chart understanding, document QA, and key information extraction. To address the loss of explicit layout analysis in end-to-end OCR, we propose Layout-as-Thought, an optional thinking phase triggered by special think tokens that generates structured layout representations -- bounding boxes, element types, and reading order -- before producing final outputs, recovering layout grounding capabilities while improving accuracy on complex layouts. Qianfan-OCR ranks first among end-to-end models on OmniDocBench v1.5 (93.12) and OlmOCR Bench (79.8), achieves competitive results on OCRBench, CCOCR, DocVQA, and ChartQA against general VLMs of comparable scale, and attains the highest average score on public key information extraction benchmarks, surpassing Gemini-3.1-Pro, Seed-2.0, and Qwen3-VL-235B. The model is publicly accessible via the Baidu AI Cloud Qianfan platform.
format Preprint
id arxiv_https___arxiv_org_abs_2603_13398
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Qianfan-OCR: A Unified End-to-End Model for Document Intelligence
Dong, Daxiang
Zheng, Mingming
Xu, Dong
Luo, Chunhua
Zhuang, Bairong
Li, Yuxuan
He, Ruoyun
Wang, Haoran
Zhang, Wenyu
Wang, Wenbo
Wang, Yicheng
Xiong, Xue
Zheng, Ayong
Zuo, Xiaoying
Ou, Ziwei
Gu, Jingnan
Guo, Quanhao
Wu, Jianmin
Yin, Dawei
Shen, Dou
Computer Vision and Pattern Recognition
We present Qianfan-OCR, a 4B-parameter end-to-end vision-language model that unifies document parsing, layout analysis, and document understanding within a single architecture. It performs direct image-to-Markdown conversion and supports diverse prompt-driven tasks including table extraction, chart understanding, document QA, and key information extraction. To address the loss of explicit layout analysis in end-to-end OCR, we propose Layout-as-Thought, an optional thinking phase triggered by special think tokens that generates structured layout representations -- bounding boxes, element types, and reading order -- before producing final outputs, recovering layout grounding capabilities while improving accuracy on complex layouts. Qianfan-OCR ranks first among end-to-end models on OmniDocBench v1.5 (93.12) and OlmOCR Bench (79.8), achieves competitive results on OCRBench, CCOCR, DocVQA, and ChartQA against general VLMs of comparable scale, and attains the highest average score on public key information extraction benchmarks, surpassing Gemini-3.1-Pro, Seed-2.0, and Qwen3-VL-235B. The model is publicly accessible via the Baidu AI Cloud Qianfan platform.
title Qianfan-OCR: A Unified End-to-End Model for Document Intelligence
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.13398