OmniDocLayout: Towards Diverse Document Layout Generation via Coarse-to-Fine LLM Learning

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Kang, Hengrui, Gu, Zhuangcheng, Zhao, Zhiyuan, Wen, Zichen, Wang, Bin, Li, Weijia, He, Conghui
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912725913829376
author Kang, Hengrui
Gu, Zhuangcheng
Zhao, Zhiyuan
Wen, Zichen
Wang, Bin
Li, Weijia
He, Conghui
author_facet Kang, Hengrui
Gu, Zhuangcheng
Zhao, Zhiyuan
Wen, Zichen
Wang, Bin
Li, Weijia
He, Conghui
contents Document AI has advanced rapidly and is attracting increasing attention. Yet, while most efforts have focused on document layout analysis (DLA), its generative counterpart, layout generation, remains underexplored. Distinct from traditional graphic layout design and room layout planning, document layout generation typically involves a larger number of elements per page and exhibits greater structural diversity and complexity. Currently, a major obstacle lies in the scarcity of diverse document layouts: academic papers with Manhattan-style structures dominate existing studies, while open-world genres such as newspapers and magazines remain severely underrepresented. To address this gap, we curate OmniDocLayout-1M, the first million-scale dataset of diverse document layouts, covering six common document types and comprising contemporary layouts collected from multiple sources. Moreover, since existing methods struggle in complex domains and often fail to arrange long sequences coherently, we introduce OmniDocLayout-LLM, a 0.5B model with designed two-stage Coarse-to-Fine learning paradigm:1) learning universal layout principles from our dataset with coarse category definitions, and 2) transferring the knowledge to a specific domain with few fine-grained annotated samples. Extensive experiments demonstrate that our approach achieves strong performance on multiple domains in M$^6$Doc dataset, substantially surpassing both existing layout generation experts and several latest general-purpose LLMs. Our code, dataset, and models will be publicly released.
format Preprint
id arxiv_https___arxiv_org_abs_2510_26213
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle OmniDocLayout: Towards Diverse Document Layout Generation via Coarse-to-Fine LLM Learning
Kang, Hengrui
Gu, Zhuangcheng
Zhao, Zhiyuan
Wen, Zichen
Wang, Bin
Li, Weijia
He, Conghui
Computer Vision and Pattern Recognition
Document AI has advanced rapidly and is attracting increasing attention. Yet, while most efforts have focused on document layout analysis (DLA), its generative counterpart, layout generation, remains underexplored. Distinct from traditional graphic layout design and room layout planning, document layout generation typically involves a larger number of elements per page and exhibits greater structural diversity and complexity. Currently, a major obstacle lies in the scarcity of diverse document layouts: academic papers with Manhattan-style structures dominate existing studies, while open-world genres such as newspapers and magazines remain severely underrepresented. To address this gap, we curate OmniDocLayout-1M, the first million-scale dataset of diverse document layouts, covering six common document types and comprising contemporary layouts collected from multiple sources. Moreover, since existing methods struggle in complex domains and often fail to arrange long sequences coherently, we introduce OmniDocLayout-LLM, a 0.5B model with designed two-stage Coarse-to-Fine learning paradigm:1) learning universal layout principles from our dataset with coarse category definitions, and 2) transferring the knowledge to a specific domain with few fine-grained annotated samples. Extensive experiments demonstrate that our approach achieves strong performance on multiple domains in M$^6$Doc dataset, substantially surpassing both existing layout generation experts and several latest general-purpose LLMs. Our code, dataset, and models will be publicly released.
title OmniDocLayout: Towards Diverse Document Layout Generation via Coarse-to-Fine LLM Learning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.26213