Docling: An Efficient Open-Source Toolkit for AI-driven Document Conversion

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Livathinos, Nikolaos, Auer, Christoph, Lysak, Maksym, Nassar, Ahmed, Dolfi, Michele, Vagenas, Panos, Ramis, Cesar Berrospi, Omenetti, Matteo, Dinkla, Kasper, Kim, Yusik, Gupta, Shubham, de Lima, Rafael Teixeira, Weber, Valery, Morin, Lucas, Meijer, Ingmar, Kuropiatnyk, Viktor, Staar, Peter W. J.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910805616754688
author Livathinos, Nikolaos
Auer, Christoph
Lysak, Maksym
Nassar, Ahmed
Dolfi, Michele
Vagenas, Panos
Ramis, Cesar Berrospi
Omenetti, Matteo
Dinkla, Kasper
Kim, Yusik
Gupta, Shubham
de Lima, Rafael Teixeira
Weber, Valery
Morin, Lucas
Meijer, Ingmar
Kuropiatnyk, Viktor
Staar, Peter W. J.
author_facet Livathinos, Nikolaos
Auer, Christoph
Lysak, Maksym
Nassar, Ahmed
Dolfi, Michele
Vagenas, Panos
Ramis, Cesar Berrospi
Omenetti, Matteo
Dinkla, Kasper
Kim, Yusik
Gupta, Shubham
de Lima, Rafael Teixeira
Weber, Valery
Morin, Lucas
Meijer, Ingmar
Kuropiatnyk, Viktor
Staar, Peter W. J.
contents We introduce Docling, an easy-to-use, self-contained, MIT-licensed, open-source toolkit for document conversion, that can parse several types of popular document formats into a unified, richly structured representation. It is powered by state-of-the-art specialized AI models for layout analysis (DocLayNet) and table structure recognition (TableFormer), and runs efficiently on commodity hardware in a small resource budget. Docling is released as a Python package and can be used as a Python API or as a CLI tool. Docling's modular architecture and efficient document representation make it easy to implement extensions, new features, models, and customizations. Docling has been already integrated in other popular open-source frameworks (e.g., LangChain, LlamaIndex, spaCy), making it a natural fit for the processing of documents and the development of high-end applications. The open-source community has fully engaged in using, promoting, and developing for Docling, which gathered 10k stars on GitHub in less than a month and was reported as the No. 1 trending repository in GitHub worldwide in November 2024.
format Preprint
id arxiv_https___arxiv_org_abs_2501_17887
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Docling: An Efficient Open-Source Toolkit for AI-driven Document Conversion
Livathinos, Nikolaos
Auer, Christoph
Lysak, Maksym
Nassar, Ahmed
Dolfi, Michele
Vagenas, Panos
Ramis, Cesar Berrospi
Omenetti, Matteo
Dinkla, Kasper
Kim, Yusik
Gupta, Shubham
de Lima, Rafael Teixeira
Weber, Valery
Morin, Lucas
Meijer, Ingmar
Kuropiatnyk, Viktor
Staar, Peter W. J.
Computation and Language
Computer Vision and Pattern Recognition
Software Engineering
We introduce Docling, an easy-to-use, self-contained, MIT-licensed, open-source toolkit for document conversion, that can parse several types of popular document formats into a unified, richly structured representation. It is powered by state-of-the-art specialized AI models for layout analysis (DocLayNet) and table structure recognition (TableFormer), and runs efficiently on commodity hardware in a small resource budget. Docling is released as a Python package and can be used as a Python API or as a CLI tool. Docling's modular architecture and efficient document representation make it easy to implement extensions, new features, models, and customizations. Docling has been already integrated in other popular open-source frameworks (e.g., LangChain, LlamaIndex, spaCy), making it a natural fit for the processing of documents and the development of high-end applications. The open-source community has fully engaged in using, promoting, and developing for Docling, which gathered 10k stars on GitHub in less than a month and was reported as the No. 1 trending repository in GitHub worldwide in November 2024.
title Docling: An Efficient Open-Source Toolkit for AI-driven Document Conversion
topic Computation and Language
Computer Vision and Pattern Recognition
Software Engineering
url https://arxiv.org/abs/2501.17887