CC-OCR V2: Benchmarking Large Multimodal Models for Literacy in Real-world Document Processing

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Xu, Zhipeng, Ji, Junhao, Chen, Zulong, Liu, Zhenghao, Liu, Qing, Peng, Chunyi, Qin, Zubao, Xu, Ze, Wan, Jianqiang, Tang, Jun, Yang, Zhibo, Bai, Shuai, Liu, Dayiheng
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915981296664576
author Xu, Zhipeng
Ji, Junhao
Chen, Zulong
Liu, Zhenghao
Liu, Qing
Peng, Chunyi
Qin, Zubao
Xu, Ze
Wan, Jianqiang
Tang, Jun
Yang, Zhibo
Bai, Shuai
Liu, Dayiheng
author_facet Xu, Zhipeng
Ji, Junhao
Chen, Zulong
Liu, Zhenghao
Liu, Qing
Peng, Chunyi
Qin, Zubao
Xu, Ze
Wan, Jianqiang
Tang, Jun
Yang, Zhibo
Bai, Shuai
Liu, Dayiheng
contents Large Multimodal Models (LMMs) have recently shown strong performance on Optical Character Recognition (OCR) tasks, demonstrating their promising capability in document literacy. However, their effectiveness in real-world applications remains underexplored, as existing benchmarks adopt task scopes misaligned with practical applications and assume homogeneous acquisition conditions. To address this gap, we introduce CC-OCR V2, a comprehensive and challenging OCR benchmark tailored to real-world document processing. CC-OCR V2 focuses on practical enterprise document processing tasks and incorporates hard and corner cases that are critical yet underrepresented in prior benchmarks, covering 5 major OCR-centric tracks: text recognition, document parsing, document grounding, key information extraction, and document question answering, comprising 7,093 high-difficulty samples. Extensive experiments on 14 advanced LMMs reveal that current models fall short of real-world application requirements. Even state-of-the-art LMMs exhibit substantial performance degradation across diverse tasks and scenarios. These findings reveal a significant gap between performance on current benchmarks and effectiveness in real-world applications. We release the full dataset and evaluation toolkit at https://github.com/eioss/CC-OCR-V2.
format Preprint
id arxiv_https___arxiv_org_abs_2605_03903
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle CC-OCR V2: Benchmarking Large Multimodal Models for Literacy in Real-world Document Processing
Xu, Zhipeng
Ji, Junhao
Chen, Zulong
Liu, Zhenghao
Liu, Qing
Peng, Chunyi
Qin, Zubao
Xu, Ze
Wan, Jianqiang
Tang, Jun
Yang, Zhibo
Bai, Shuai
Liu, Dayiheng
Computation and Language
Large Multimodal Models (LMMs) have recently shown strong performance on Optical Character Recognition (OCR) tasks, demonstrating their promising capability in document literacy. However, their effectiveness in real-world applications remains underexplored, as existing benchmarks adopt task scopes misaligned with practical applications and assume homogeneous acquisition conditions. To address this gap, we introduce CC-OCR V2, a comprehensive and challenging OCR benchmark tailored to real-world document processing. CC-OCR V2 focuses on practical enterprise document processing tasks and incorporates hard and corner cases that are critical yet underrepresented in prior benchmarks, covering 5 major OCR-centric tracks: text recognition, document parsing, document grounding, key information extraction, and document question answering, comprising 7,093 high-difficulty samples. Extensive experiments on 14 advanced LMMs reveal that current models fall short of real-world application requirements. Even state-of-the-art LMMs exhibit substantial performance degradation across diverse tasks and scenarios. These findings reveal a significant gap between performance on current benchmarks and effectiveness in real-world applications. We release the full dataset and evaluation toolkit at https://github.com/eioss/CC-OCR-V2.
title CC-OCR V2: Benchmarking Large Multimodal Models for Literacy in Real-world Document Processing
topic Computation and Language
url https://arxiv.org/abs/2605.03903