Advanced Layout Analysis Models for Docling

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Livathinos, Nikolaos, Auer, Christoph, Nassar, Ahmed, de Lima, Rafael Teixeira, Lysak, Maksym, Ebouky, Brown, Berrospi, Cesar, Dolfi, Michele, Vagenas, Panagiotis, Omenetti, Matteo, Dinkla, Kasper, Kim, Yusik, Weber, Valery, Morin, Lucas, Meijer, Ingmar, Kuropiatnyk, Viktor, Strohmeyer, Tim, Gurbuz, A. Said, Staar, Peter W. J.
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912587556323328
author Livathinos, Nikolaos
Auer, Christoph
Nassar, Ahmed
de Lima, Rafael Teixeira
Lysak, Maksym
Ebouky, Brown
Berrospi, Cesar
Dolfi, Michele
Vagenas, Panagiotis
Omenetti, Matteo
Dinkla, Kasper
Kim, Yusik
Weber, Valery
Morin, Lucas
Meijer, Ingmar
Kuropiatnyk, Viktor
Strohmeyer, Tim
Gurbuz, A. Said
Staar, Peter W. J.
author_facet Livathinos, Nikolaos
Auer, Christoph
Nassar, Ahmed
de Lima, Rafael Teixeira
Lysak, Maksym
Ebouky, Brown
Berrospi, Cesar
Dolfi, Michele
Vagenas, Panagiotis
Omenetti, Matteo
Dinkla, Kasper
Kim, Yusik
Weber, Valery
Morin, Lucas
Meijer, Ingmar
Kuropiatnyk, Viktor
Strohmeyer, Tim
Gurbuz, A. Said
Staar, Peter W. J.
contents This technical report documents the development of novel Layout Analysis models integrated into the Docling document-conversion pipeline. We trained several state-of-the-art object detectors based on the RT-DETR, RT-DETRv2 and DFINE architectures on a heterogeneous corpus of 150,000 documents (both openly available and proprietary). Post-processing steps were applied to the raw detections to make them more applicable to the document conversion task. We evaluated the effectiveness of the layout analysis on various document benchmarks using different methodologies while also measuring the runtime performance across different environments (CPU, Nvidia and Apple GPUs). We introduce five new document layout models achieving 20.6% - 23.9% mAP improvement over Docling's previous baseline, with comparable or better runtime. Our best model, "heron-101", attains 78% mAP with 28 ms/image inference time on a single NVIDIA A100 GPU. Extensive quantitative and qualitative experiments establish best practices for training, evaluating, and deploying document-layout detectors, providing actionable guidance for the document conversion community. All trained checkpoints, code, and documentation are released under a permissive license on HuggingFace.
format Preprint
id arxiv_https___arxiv_org_abs_2509_11720
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Advanced Layout Analysis Models for Docling
Livathinos, Nikolaos
Auer, Christoph
Nassar, Ahmed
de Lima, Rafael Teixeira
Lysak, Maksym
Ebouky, Brown
Berrospi, Cesar
Dolfi, Michele
Vagenas, Panagiotis
Omenetti, Matteo
Dinkla, Kasper
Kim, Yusik
Weber, Valery
Morin, Lucas
Meijer, Ingmar
Kuropiatnyk, Viktor
Strohmeyer, Tim
Gurbuz, A. Said
Staar, Peter W. J.
Computer Vision and Pattern Recognition
This technical report documents the development of novel Layout Analysis models integrated into the Docling document-conversion pipeline. We trained several state-of-the-art object detectors based on the RT-DETR, RT-DETRv2 and DFINE architectures on a heterogeneous corpus of 150,000 documents (both openly available and proprietary). Post-processing steps were applied to the raw detections to make them more applicable to the document conversion task. We evaluated the effectiveness of the layout analysis on various document benchmarks using different methodologies while also measuring the runtime performance across different environments (CPU, Nvidia and Apple GPUs). We introduce five new document layout models achieving 20.6% - 23.9% mAP improvement over Docling's previous baseline, with comparable or better runtime. Our best model, "heron-101", attains 78% mAP with 28 ms/image inference time on a single NVIDIA A100 GPU. Extensive quantitative and qualitative experiments establish best practices for training, evaluating, and deploying document-layout detectors, providing actionable guidance for the document conversion community. All trained checkpoints, code, and documentation are released under a permissive license on HuggingFace.
title Advanced Layout Analysis Models for Docling
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.11720