DocAtlas: Multilingual Document Understanding Across 80+ Languages

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Heakl, Ahmed, Mohamed, Youssef, Sohail, Abdullah, Elbadry, Rania, Nassar, Ahmed, Staar, Peter W. J., Khan, Fahad Shahbaz, Razzak, Imran, Khan, Salman
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914586788102144
author Heakl, Ahmed
Mohamed, Youssef
Sohail, Abdullah
Elbadry, Rania
Nassar, Ahmed
Staar, Peter W. J.
Khan, Fahad Shahbaz
Razzak, Imran
Khan, Salman
author_facet Heakl, Ahmed
Mohamed, Youssef
Sohail, Abdullah
Elbadry, Rania
Nassar, Ahmed
Staar, Peter W. J.
Khan, Fahad Shahbaz
Razzak, Imran
Khan, Salman
contents Multilingual document understanding remains limited for low-resource languages due to scarce training data and model-based annotation pipelines that perpetuate existing biases. We introduce DocAtlas, a framework that constructs high-fidelity OCR datasets and benchmarks covering 82 languages and 9 evaluation tasks. Our dual pipelines, differential rendering of native DOCX documents and synthetic LaTeX-based generation for right-to-left scripts produce precise structural annotations in a unified DocTag format encoding layout, text, and component types, without learned models for core annotation. Evaluating 16 state-of-the-art models reveals persistent gaps in low-resource scripts. We show that Direct Preference Optimization (DPO) using rendering-derived ground truth as positive signal achieves stable multilingual adaptation, improving both in-domain (+1.9%) and out-of-domain (+1.8%) accuracy without measurable base-language degradation, where supervised fine-tuning degrades out-of-domain performance by up to 21%. Our best variant, DocAtlas-DeepSeek, improves +1.7% over the strongest baseline. Code is available at https://github.com/ahmedheakl/DocAtlas .
format Preprint
id arxiv_https___arxiv_org_abs_2605_12623
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle DocAtlas: Multilingual Document Understanding Across 80+ Languages
Heakl, Ahmed
Mohamed, Youssef
Sohail, Abdullah
Elbadry, Rania
Nassar, Ahmed
Staar, Peter W. J.
Khan, Fahad Shahbaz
Razzak, Imran
Khan, Salman
Computation and Language
Computer Vision and Pattern Recognition
Machine Learning
Multilingual document understanding remains limited for low-resource languages due to scarce training data and model-based annotation pipelines that perpetuate existing biases. We introduce DocAtlas, a framework that constructs high-fidelity OCR datasets and benchmarks covering 82 languages and 9 evaluation tasks. Our dual pipelines, differential rendering of native DOCX documents and synthetic LaTeX-based generation for right-to-left scripts produce precise structural annotations in a unified DocTag format encoding layout, text, and component types, without learned models for core annotation. Evaluating 16 state-of-the-art models reveals persistent gaps in low-resource scripts. We show that Direct Preference Optimization (DPO) using rendering-derived ground truth as positive signal achieves stable multilingual adaptation, improving both in-domain (+1.9%) and out-of-domain (+1.8%) accuracy without measurable base-language degradation, where supervised fine-tuning degrades out-of-domain performance by up to 21%. Our best variant, DocAtlas-DeepSeek, improves +1.7% over the strongest baseline. Code is available at https://github.com/ahmedheakl/DocAtlas .
title DocAtlas: Multilingual Document Understanding Across 80+ Languages
topic Computation and Language
Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2605.12623