DPAR: Dynamic Patchification for Efficient Autoregressive Visual Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Srivastava, Divyansh, Mehra, Akshay, Maneriker, Pranav, Sanyal, Debopam, Raj, Vishnu, Kamarshi, Vijay, Du, Fan, Kimball, Joshua
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911339939627008
author Srivastava, Divyansh
Mehra, Akshay
Maneriker, Pranav
Sanyal, Debopam
Raj, Vishnu
Kamarshi, Vijay
Du, Fan
Kimball, Joshua
author_facet Srivastava, Divyansh
Mehra, Akshay
Maneriker, Pranav
Sanyal, Debopam
Raj, Vishnu
Kamarshi, Vijay
Du, Fan
Kimball, Joshua
contents Decoder-only autoregressive image generation typically relies on fixed-length tokenization schemes whose token counts grow quadratically with resolution, substantially increasing the computational and memory demands of attention. We present DPAR, a novel decoder-only autoregressive model that dynamically aggregates image tokens into a variable number of patches for efficient image generation. Our work is the first to demonstrate that next-token prediction entropy from a lightweight and unsupervised autoregressive model provides a reliable criterion for merging tokens into larger patches based on information content. DPAR makes minimal modifications to the standard decoder architecture, ensuring compatibility with multimodal generation frameworks and allocating more compute to generation of high-information image regions. Further, we demonstrate that training with dynamically sized patches yields representations that are robust to patch boundaries, allowing DPAR to scale to larger patch sizes at inference. DPAR reduces token count by 1.81x and 2.06x on Imagenet 256 and 384 generation resolution respectively, leading to a reduction of up to 40% FLOPs in training costs. Further, our method exhibits faster convergence and improves FID by up to 27.1% relative to baseline models.
format Preprint
id arxiv_https___arxiv_org_abs_2512_21867
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DPAR: Dynamic Patchification for Efficient Autoregressive Visual Generation
Srivastava, Divyansh
Mehra, Akshay
Maneriker, Pranav
Sanyal, Debopam
Raj, Vishnu
Kamarshi, Vijay
Du, Fan
Kimball, Joshua
Computer Vision and Pattern Recognition
Decoder-only autoregressive image generation typically relies on fixed-length tokenization schemes whose token counts grow quadratically with resolution, substantially increasing the computational and memory demands of attention. We present DPAR, a novel decoder-only autoregressive model that dynamically aggregates image tokens into a variable number of patches for efficient image generation. Our work is the first to demonstrate that next-token prediction entropy from a lightweight and unsupervised autoregressive model provides a reliable criterion for merging tokens into larger patches based on information content. DPAR makes minimal modifications to the standard decoder architecture, ensuring compatibility with multimodal generation frameworks and allocating more compute to generation of high-information image regions. Further, we demonstrate that training with dynamically sized patches yields representations that are robust to patch boundaries, allowing DPAR to scale to larger patch sizes at inference. DPAR reduces token count by 1.81x and 2.06x on Imagenet 256 and 384 generation resolution respectively, leading to a reduction of up to 40% FLOPs in training costs. Further, our method exhibits faster convergence and improves FID by up to 27.1% relative to baseline models.
title DPAR: Dynamic Patchification for Efficient Autoregressive Visual Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.21867