HybriDLA: Hybrid Generation for Document Layout Analysis

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Yufan, Moured, Omar, Liu, Ruiping, Zheng, Junwei, Peng, Kunyu, Zhang, Jiaming, Stiefelhagen, Rainer
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912727602036736
author Chen, Yufan
Moured, Omar
Liu, Ruiping
Zheng, Junwei
Peng, Kunyu
Zhang, Jiaming
Stiefelhagen, Rainer
author_facet Chen, Yufan
Moured, Omar
Liu, Ruiping
Zheng, Junwei
Peng, Kunyu
Zhang, Jiaming
Stiefelhagen, Rainer
contents Conventional document layout analysis (DLA) traditionally depends on empirical priors or a fixed set of learnable queries executed in a single forward pass. While sufficient for early-generation documents with a small, predetermined number of regions, this paradigm struggles with contemporary documents, which exhibit diverse element counts and increasingly complex layouts. To address challenges posed by modern documents, we present HybriDLA, a novel generative framework that unifies diffusion and autoregressive decoding within a single layer. The diffusion component iteratively refines bounding-box hypotheses, whereas the autoregressive component injects semantic and contextual awareness, enabling precise region prediction even in highly varied layouts. To further enhance detection quality, we design a multi-scale feature-fusion encoder that captures both fine-grained and high-level visual cues. This architecture elevates performance to 83.5% mean Average Precision (mAP). Extensive experiments on the DocLayNet and M$^6$Doc benchmarks demonstrate that HybriDLA sets a state-of-the-art performance, outperforming previous approaches. All data and models will be made publicly available at https://yufanchen96.github.io/projects/HybriDLA.
format Preprint
id arxiv_https___arxiv_org_abs_2511_19919
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle HybriDLA: Hybrid Generation for Document Layout Analysis
Chen, Yufan
Moured, Omar
Liu, Ruiping
Zheng, Junwei
Peng, Kunyu
Zhang, Jiaming
Stiefelhagen, Rainer
Computer Vision and Pattern Recognition
Conventional document layout analysis (DLA) traditionally depends on empirical priors or a fixed set of learnable queries executed in a single forward pass. While sufficient for early-generation documents with a small, predetermined number of regions, this paradigm struggles with contemporary documents, which exhibit diverse element counts and increasingly complex layouts. To address challenges posed by modern documents, we present HybriDLA, a novel generative framework that unifies diffusion and autoregressive decoding within a single layer. The diffusion component iteratively refines bounding-box hypotheses, whereas the autoregressive component injects semantic and contextual awareness, enabling precise region prediction even in highly varied layouts. To further enhance detection quality, we design a multi-scale feature-fusion encoder that captures both fine-grained and high-level visual cues. This architecture elevates performance to 83.5% mean Average Precision (mAP). Extensive experiments on the DocLayNet and M$^6$Doc benchmarks demonstrate that HybriDLA sets a state-of-the-art performance, outperforming previous approaches. All data and models will be made publicly available at https://yufanchen96.github.io/projects/HybriDLA.
title HybriDLA: Hybrid Generation for Document Layout Analysis
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.19919