Prefix-Adaptive Block Diffusion for Efficient Document Recognition

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Chai, Mingxu, Shen, Ziyu, Liu, Chenyu, Zhang, Kaidi, Zhang, Jiazheng, Zhu, Dingwei, Xi, Zhiheng, Chen, Ruoyu, Long, Jun, Kang, Jihua, Gui, Tao, Zhang, Qi
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918506311712768
author Chai, Mingxu
Shen, Ziyu
Liu, Chenyu
Zhang, Kaidi
Zhang, Jiazheng
Zhu, Dingwei
Xi, Zhiheng
Chen, Ruoyu
Long, Jun
Kang, Jihua
Gui, Tao
Zhang, Qi
author_facet Chai, Mingxu
Shen, Ziyu
Liu, Chenyu
Zhang, Kaidi
Zhang, Jiazheng
Zhu, Dingwei
Xi, Zhiheng
Chen, Ruoyu
Long, Jun
Kang, Jihua
Gui, Tao
Zhang, Qi
contents Block Diffusion Models (BDMs) support parallel generation, flexible-length output, and KV caching, making them promising for efficient document parsing. However, existing BDMs bind denoising and cache commitment to fixed block boundaries: parallelism shrinks during intra-block denoising, while generated tokens cannot be cached until the whole block is completed. Moreover, intra-block bidirectional denoising conflicts with inter-block autoregression, creating inconsistent information flow that can challenge structure-sensitive recognition. We propose the Prefix-Adaptive Block Diffusion Model (PA-BDM), which replaces intra-block bidirectional denoising with causal denoising from prefix to suffix and treats the block size as a maximum candidate range rather than a fixed commitment unit. PA-BDM uses Confidence-gated Structural Loss (CSL) to build low-entropy prefixes before extending training to longer continuations. During inference, Progressive Prefix Commitment (PPC) then dynamically commits the longest reliable prefix into the KV cache and resets the next candidate range from the updated prefix, restoring a large parallel decoding space at each step. Experiments show that the 3B PA-BDM achieves higher recognition scores on several benchmarks and improves inference throughput by 71.6\% over the 2.5B MinerU-Diffusion.
format Preprint
id arxiv_https___arxiv_org_abs_2605_16861
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Prefix-Adaptive Block Diffusion for Efficient Document Recognition
Chai, Mingxu
Shen, Ziyu
Liu, Chenyu
Zhang, Kaidi
Zhang, Jiazheng
Zhu, Dingwei
Xi, Zhiheng
Chen, Ruoyu
Long, Jun
Kang, Jihua
Gui, Tao
Zhang, Qi
Computer Vision and Pattern Recognition
Artificial Intelligence
Block Diffusion Models (BDMs) support parallel generation, flexible-length output, and KV caching, making them promising for efficient document parsing. However, existing BDMs bind denoising and cache commitment to fixed block boundaries: parallelism shrinks during intra-block denoising, while generated tokens cannot be cached until the whole block is completed. Moreover, intra-block bidirectional denoising conflicts with inter-block autoregression, creating inconsistent information flow that can challenge structure-sensitive recognition. We propose the Prefix-Adaptive Block Diffusion Model (PA-BDM), which replaces intra-block bidirectional denoising with causal denoising from prefix to suffix and treats the block size as a maximum candidate range rather than a fixed commitment unit. PA-BDM uses Confidence-gated Structural Loss (CSL) to build low-entropy prefixes before extending training to longer continuations. During inference, Progressive Prefix Commitment (PPC) then dynamically commits the longest reliable prefix into the KV cache and resets the next candidate range from the updated prefix, restoring a large parallel decoding space at each step. Experiments show that the 3B PA-BDM achieves higher recognition scores on several benchmarks and improves inference throughput by 71.6\% over the 2.5B MinerU-Diffusion.
title Prefix-Adaptive Block Diffusion for Efficient Document Recognition
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2605.16861