olmOCR 2: Unit Test Rewards for Document OCR

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Poznanski, Jake, Soldaini, Luca, Lo, Kyle
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908606320869376
author Poznanski, Jake
Soldaini, Luca
Lo, Kyle
author_facet Poznanski, Jake
Soldaini, Luca
Lo, Kyle
contents We present olmOCR 2, the latest in our family of powerful OCR systems for converting digitized print documents, like PDFs, into clean, naturally ordered plain text. olmOCR 2 is powered by olmOCR-2-7B-1025, a specialized, 7B vision language model (VLM) trained using reinforcement learning with verifiable rewards (RLVR), where our rewards are a diverse set of binary unit tests. To scale unit test creation, we develop a pipeline for generating synthetic documents with diverse and challenging layouts, known ground-truth HTML source code, and extracted test cases. We show that RL training on these test cases results in state-of-the-art performance on olmOCR-Bench, our English-language OCR benchmark, with the largest improvements in math formula conversion, table parsing, and multi-column layouts compared to previous versions. We release our model, data and code under permissive open licenses.
format Preprint
id arxiv_https___arxiv_org_abs_2510_19817
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle olmOCR 2: Unit Test Rewards for Document OCR
Poznanski, Jake
Soldaini, Luca
Lo, Kyle
Computer Vision and Pattern Recognition
Computation and Language
We present olmOCR 2, the latest in our family of powerful OCR systems for converting digitized print documents, like PDFs, into clean, naturally ordered plain text. olmOCR 2 is powered by olmOCR-2-7B-1025, a specialized, 7B vision language model (VLM) trained using reinforcement learning with verifiable rewards (RLVR), where our rewards are a diverse set of binary unit tests. To scale unit test creation, we develop a pipeline for generating synthetic documents with diverse and challenging layouts, known ground-truth HTML source code, and extracted test cases. We show that RL training on these test cases results in state-of-the-art performance on olmOCR-Bench, our English-language OCR benchmark, with the largest improvements in math formula conversion, table parsing, and multi-column layouts compared to previous versions. We release our model, data and code under permissive open licenses.
title olmOCR 2: Unit Test Rewards for Document OCR
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2510.19817