DocAtlas: Multilingual Document Understanding Across 80+ Languages
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866914586788102144 |
|---|---|
| author | Heakl, Ahmed Mohamed, Youssef Sohail, Abdullah Elbadry, Rania Nassar, Ahmed Staar, Peter W. J. Khan, Fahad Shahbaz Razzak, Imran Khan, Salman |
| author_facet | Heakl, Ahmed Mohamed, Youssef Sohail, Abdullah Elbadry, Rania Nassar, Ahmed Staar, Peter W. J. Khan, Fahad Shahbaz Razzak, Imran Khan, Salman |
| contents | Multilingual document understanding remains limited for low-resource languages due to scarce training data and model-based annotation pipelines that perpetuate existing biases. We introduce DocAtlas, a framework that constructs high-fidelity OCR datasets and benchmarks covering 82 languages and 9 evaluation tasks. Our dual pipelines, differential rendering of native DOCX documents and synthetic LaTeX-based generation for right-to-left scripts produce precise structural annotations in a unified DocTag format encoding layout, text, and component types, without learned models for core annotation. Evaluating 16 state-of-the-art models reveals persistent gaps in low-resource scripts. We show that Direct Preference Optimization (DPO) using rendering-derived ground truth as positive signal achieves stable multilingual adaptation, improving both in-domain (+1.9%) and out-of-domain (+1.8%) accuracy without measurable base-language degradation, where supervised fine-tuning degrades out-of-domain performance by up to 21%. Our best variant, DocAtlas-DeepSeek, improves +1.7% over the strongest baseline. Code is available at https://github.com/ahmedheakl/DocAtlas . |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2605_12623 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | DocAtlas: Multilingual Document Understanding Across 80+ Languages Heakl, Ahmed Mohamed, Youssef Sohail, Abdullah Elbadry, Rania Nassar, Ahmed Staar, Peter W. J. Khan, Fahad Shahbaz Razzak, Imran Khan, Salman Computation and Language Computer Vision and Pattern Recognition Machine Learning Multilingual document understanding remains limited for low-resource languages due to scarce training data and model-based annotation pipelines that perpetuate existing biases. We introduce DocAtlas, a framework that constructs high-fidelity OCR datasets and benchmarks covering 82 languages and 9 evaluation tasks. Our dual pipelines, differential rendering of native DOCX documents and synthetic LaTeX-based generation for right-to-left scripts produce precise structural annotations in a unified DocTag format encoding layout, text, and component types, without learned models for core annotation. Evaluating 16 state-of-the-art models reveals persistent gaps in low-resource scripts. We show that Direct Preference Optimization (DPO) using rendering-derived ground truth as positive signal achieves stable multilingual adaptation, improving both in-domain (+1.9%) and out-of-domain (+1.8%) accuracy without measurable base-language degradation, where supervised fine-tuning degrades out-of-domain performance by up to 21%. Our best variant, DocAtlas-DeepSeek, improves +1.7% over the strongest baseline. Code is available at https://github.com/ahmedheakl/DocAtlas . |
| title | DocAtlas: Multilingual Document Understanding Across 80+ Languages |
| topic | Computation and Language Computer Vision and Pattern Recognition Machine Learning |
| url | https://arxiv.org/abs/2605.12623 |