Neighboring Autoregressive Modeling for Efficient Visual Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: He, Yefei, He, Yuanyu, He, Shaoxuan, Chen, Feng, Zhou, Hong, Zhang, Kaipeng, Zhuang, Bohan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916652281495552
author He, Yefei
He, Yuanyu
He, Shaoxuan
Chen, Feng
Zhou, Hong
Zhang, Kaipeng
Zhuang, Bohan
author_facet He, Yefei
He, Yuanyu
He, Shaoxuan
Chen, Feng
Zhou, Hong
Zhang, Kaipeng
Zhuang, Bohan
contents Visual autoregressive models typically adhere to a raster-order ``next-token prediction" paradigm, which overlooks the spatial and temporal locality inherent in visual content. Specifically, visual tokens exhibit significantly stronger correlations with their spatially or temporally adjacent tokens compared to those that are distant. In this paper, we propose Neighboring Autoregressive Modeling (NAR), a novel paradigm that formulates autoregressive visual generation as a progressive outpainting procedure, following a near-to-far ``next-neighbor prediction" mechanism. Starting from an initial token, the remaining tokens are decoded in ascending order of their Manhattan distance from the initial token in the spatial-temporal space, progressively expanding the boundary of the decoded region. To enable parallel prediction of multiple adjacent tokens in the spatial-temporal space, we introduce a set of dimension-oriented decoding heads, each predicting the next token along a mutually orthogonal dimension. During inference, all tokens adjacent to the decoded tokens are processed in parallel, substantially reducing the model forward steps for generation. Experiments on ImageNet$256\times 256$ and UCF101 demonstrate that NAR achieves 2.4$\times$ and 8.6$\times$ higher throughput respectively, while obtaining superior FID/FVD scores for both image and video generation tasks compared to the PAR-4X approach. When evaluating on text-to-image generation benchmark GenEval, NAR with 0.8B parameters outperforms Chameleon-7B while using merely 0.4 of the training data. Code is available at https://github.com/ThisisBillhe/NAR.
format Preprint
id arxiv_https___arxiv_org_abs_2503_10696
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Neighboring Autoregressive Modeling for Efficient Visual Generation
He, Yefei
He, Yuanyu
He, Shaoxuan
Chen, Feng
Zhou, Hong
Zhang, Kaipeng
Zhuang, Bohan
Computer Vision and Pattern Recognition
Image and Video Processing
Visual autoregressive models typically adhere to a raster-order ``next-token prediction" paradigm, which overlooks the spatial and temporal locality inherent in visual content. Specifically, visual tokens exhibit significantly stronger correlations with their spatially or temporally adjacent tokens compared to those that are distant. In this paper, we propose Neighboring Autoregressive Modeling (NAR), a novel paradigm that formulates autoregressive visual generation as a progressive outpainting procedure, following a near-to-far ``next-neighbor prediction" mechanism. Starting from an initial token, the remaining tokens are decoded in ascending order of their Manhattan distance from the initial token in the spatial-temporal space, progressively expanding the boundary of the decoded region. To enable parallel prediction of multiple adjacent tokens in the spatial-temporal space, we introduce a set of dimension-oriented decoding heads, each predicting the next token along a mutually orthogonal dimension. During inference, all tokens adjacent to the decoded tokens are processed in parallel, substantially reducing the model forward steps for generation. Experiments on ImageNet$256\times 256$ and UCF101 demonstrate that NAR achieves 2.4$\times$ and 8.6$\times$ higher throughput respectively, while obtaining superior FID/FVD scores for both image and video generation tasks compared to the PAR-4X approach. When evaluating on text-to-image generation benchmark GenEval, NAR with 0.8B parameters outperforms Chameleon-7B while using merely 0.4 of the training data. Code is available at https://github.com/ThisisBillhe/NAR.
title Neighboring Autoregressive Modeling for Efficient Visual Generation
topic Computer Vision and Pattern Recognition
Image and Video Processing
url https://arxiv.org/abs/2503.10696