Beyond Next-Token: Next-X Prediction for Autoregressive Visual Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ren, Sucheng, Yu, Qihang, He, Ju, Shen, Xiaohui, Yuille, Alan, Chen, Liang-Chieh |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
FlowAR: Scale-wise Autoregressive Image Generation Meets Flow Matching
von: Ren, Sucheng, et al.
Veröffentlicht: (2024)
von: Ren, Sucheng, et al.
Veröffentlicht: (2024)
Frequency-Aware Flow Matching for High-Quality Image Generation
von: Ren, Sucheng, et al.
Veröffentlicht: (2026)
von: Ren, Sucheng, et al.
Veröffentlicht: (2026)
Grouping First, Attending Smartly: Training-Free Acceleration for Diffusion Transformers
von: Ren, Sucheng, et al.
Veröffentlicht: (2025)
von: Ren, Sucheng, et al.
Veröffentlicht: (2025)
Autoregressive Video Generation beyond Next Frames Prediction
von: Ren, Sucheng, et al.
Veröffentlicht: (2025)
von: Ren, Sucheng, et al.
Veröffentlicht: (2025)
Randomized Autoregressive Visual Generation
von: Yu, Qihang, et al.
Veröffentlicht: (2024)
von: Yu, Qihang, et al.
Veröffentlicht: (2024)
ViTamin: Designing Scalable Vision Models in the Vision-Language Era
von: Chen, Jieneng, et al.
Veröffentlicht: (2024)
von: Chen, Jieneng, et al.
Veröffentlicht: (2024)
A Simple Video Segmenter by Tracking Objects Along Axial Trajectories
von: He, Ju, et al.
Veröffentlicht: (2023)
von: He, Ju, et al.
Veröffentlicht: (2023)
ReVision: Refining Video Diffusion with Explicit 3D Motion Modeling
von: Liu, Qihao, et al.
Veröffentlicht: (2025)
von: Liu, Qihao, et al.
Veröffentlicht: (2025)
M-VAR: Decoupled Scale-wise Autoregressive Modeling for High-Quality Image Generation
von: Ren, Sucheng, et al.
Veröffentlicht: (2024)
von: Ren, Sucheng, et al.
Veröffentlicht: (2024)
Large Language Models are Universal Reasoners for Visual Generation
von: Ren, Sucheng, et al.
Veröffentlicht: (2026)
von: Ren, Sucheng, et al.
Veröffentlicht: (2026)
Democratizing Text-to-Image Masked Generative Models with Compact Text-Aware One-Dimensional Tokens
von: Kim, Dongwon, et al.
Veröffentlicht: (2025)
von: Kim, Dongwon, et al.
Veröffentlicht: (2025)
ARVideo: Autoregressive Pretraining for Self-Supervised Video Representation Learning
von: Ren, Sucheng, et al.
Veröffentlicht: (2024)
von: Ren, Sucheng, et al.
Veröffentlicht: (2024)
Next Patch Prediction for Autoregressive Visual Generation
von: Pang, Yatian, et al.
Veröffentlicht: (2024)
von: Pang, Yatian, et al.
Veröffentlicht: (2024)
An Image is Worth 32 Tokens for Reconstruction and Generation
von: Yu, Qihang, et al.
Veröffentlicht: (2024)
von: Yu, Qihang, et al.
Veröffentlicht: (2024)
FlowTok: Flowing Seamlessly Across Text and Image Tokens
von: He, Ju, et al.
Veröffentlicht: (2025)
von: He, Ju, et al.
Veröffentlicht: (2025)
Autoregressive Image Generation with Masked Bit Modeling
von: Yu, Qihang, et al.
Veröffentlicht: (2026)
von: Yu, Qihang, et al.
Veröffentlicht: (2026)
Dictionary-based Framework for Interpretable and Consistent Object Parsing
von: Zhang, Tiezheng, et al.
Veröffentlicht: (2025)
von: Zhang, Tiezheng, et al.
Veröffentlicht: (2025)
Alleviating Distortion in Image Generation via Multi-Resolution Diffusion Models and Time-Dependent Layer Normalization
von: Liu, Qihao, et al.
Veröffentlicht: (2024)
von: Liu, Qihao, et al.
Veröffentlicht: (2024)
FVAR: Visual Autoregressive Modeling via Next Focus Prediction
von: Li, Xiaofan, et al.
Veröffentlicht: (2025)
von: Li, Xiaofan, et al.
Veröffentlicht: (2025)
ViMix-14M: A Curated Multi-Source Video-Text Dataset with Long-Form, High-Quality Captions and Crawl-Free Access
von: Yang, Timing, et al.
Veröffentlicht: (2025)
von: Yang, Timing, et al.
Veröffentlicht: (2025)
AMP: Autoregressive Motion Prediction Revisited with Next Token Prediction for Autonomous Driving
von: Jia, Xiaosong, et al.
Veröffentlicht: (2024)
von: Jia, Xiaosong, et al.
Veröffentlicht: (2024)
Rejuvenating image-GPT as Strong Visual Representation Learners
von: Ren, Sucheng, et al.
Veröffentlicht: (2023)
von: Ren, Sucheng, et al.
Veröffentlicht: (2023)
NEP: Autoregressive Image Editing via Next Editing Token Prediction
von: Wu, Huimin, et al.
Veröffentlicht: (2025)
von: Wu, Huimin, et al.
Veröffentlicht: (2025)
MaskBit: Embedding-free Image Generation via Bit Tokens
von: Weber, Mark, et al.
Veröffentlicht: (2024)
von: Weber, Mark, et al.
Veröffentlicht: (2024)
NextStep-1: Toward Autoregressive Image Generation with Continuous Tokens at Scale
von: NextStep Team, et al.
Veröffentlicht: (2025)
von: NextStep Team, et al.
Veröffentlicht: (2025)
A Frame is Worth One Token: Efficient Generative World Modeling with Delta Tokens
von: Kerssies, Tommie, et al.
Veröffentlicht: (2026)
von: Kerssies, Tommie, et al.
Veröffentlicht: (2026)
ARMesh: Autoregressive Mesh Generation via Next-Level-of-Detail Prediction
von: Lei, Jiabao, et al.
Veröffentlicht: (2025)
von: Lei, Jiabao, et al.
Veröffentlicht: (2025)
Object Recognition as Next Token Prediction
von: Yue, Kaiyu, et al.
Veröffentlicht: (2023)
von: Yue, Kaiyu, et al.
Veröffentlicht: (2023)
Next Block Prediction: Video Generation via Semi-Autoregressive Modeling
von: Ren, Shuhuai, et al.
Veröffentlicht: (2025)
von: Ren, Shuhuai, et al.
Veröffentlicht: (2025)
COCONut: Modernizing COCO Segmentation
von: Deng, Xueqing, et al.
Veröffentlicht: (2024)
von: Deng, Xueqing, et al.
Veröffentlicht: (2024)
Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction
von: Tian, Keyu, et al.
Veröffentlicht: (2024)
von: Tian, Keyu, et al.
Veröffentlicht: (2024)
GenAR: Next-Scale Autoregressive Generation for Spatial Gene Expression Prediction
von: Ouyang, Jiarui, et al.
Veröffentlicht: (2025)
von: Ouyang, Jiarui, et al.
Veröffentlicht: (2025)
Autoregressive Pretraining with Mamba in Vision
von: Ren, Sucheng, et al.
Veröffentlicht: (2024)
von: Ren, Sucheng, et al.
Veröffentlicht: (2024)
Arbitrary Ratio Feature Compression via Next Token Prediction
von: Liu, Yufan, et al.
Veröffentlicht: (2026)
von: Liu, Yufan, et al.
Veröffentlicht: (2026)
ScaleMoGen: Autoregressive Next-Scale Prediction for Human Motion Generation
von: Hwang, Inwoo, et al.
Veröffentlicht: (2026)
von: Hwang, Inwoo, et al.
Veröffentlicht: (2026)
Next-Scale Autoregressive Models for Text-to-Motion Generation
von: Zheng, Zhiwei, et al.
Veröffentlicht: (2026)
von: Zheng, Zhiwei, et al.
Veröffentlicht: (2026)
Long-Context Autoregressive Video Modeling with Next-Frame Prediction
von: Gu, Yuchao, et al.
Veröffentlicht: (2025)
von: Gu, Yuchao, et al.
Veröffentlicht: (2025)
TAPNext: Tracking Any Point (TAP) as Next Token Prediction
von: Zholus, Artem, et al.
Veröffentlicht: (2025)
von: Zholus, Artem, et al.
Veröffentlicht: (2025)
Emu3: Next-Token Prediction is All You Need
von: Wang, Xinlong, et al.
Veröffentlicht: (2024)
von: Wang, Xinlong, et al.
Veröffentlicht: (2024)
SPFormer: Enhancing Vision Transformer with Superpixel Representation
von: Mei, Jieru, et al.
Veröffentlicht: (2024)
von: Mei, Jieru, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
FlowAR: Scale-wise Autoregressive Image Generation Meets Flow Matching
von: Ren, Sucheng, et al.
Veröffentlicht: (2024) -
Frequency-Aware Flow Matching for High-Quality Image Generation
von: Ren, Sucheng, et al.
Veröffentlicht: (2026) -
Grouping First, Attending Smartly: Training-Free Acceleration for Diffusion Transformers
von: Ren, Sucheng, et al.
Veröffentlicht: (2025) -
Autoregressive Video Generation beyond Next Frames Prediction
von: Ren, Sucheng, et al.
Veröffentlicht: (2025) -
Randomized Autoregressive Visual Generation
von: Yu, Qihang, et al.
Veröffentlicht: (2024)