Autoregressive Image Generation Needs Only a Few Lines of Cached Tokens
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Qin, Ziran, Lv, Youru, Lin, Mingbao, Zhang, Zeren, Gan, Chanfan, Chen, Tieyuan, Lin, Weiyao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Head-Aware KV Cache Compression for Efficient Visual Autoregressive Modeling
von: Qin, Ziran, et al.
Veröffentlicht: (2025)
von: Qin, Ziran, et al.
Veröffentlicht: (2025)
Massive Activations are the Key to Local Detail Synthesis in Diffusion Transformers
von: Gan, Chaofan, et al.
Veröffentlicht: (2025)
von: Gan, Chaofan, et al.
Veröffentlicht: (2025)
VidLaDA: Bidirectional Diffusion Large Language Models for Efficient Video Understanding
von: He, Zhihao, et al.
Veröffentlicht: (2026)
von: He, Zhihao, et al.
Veröffentlicht: (2026)
Looking Beyond Visible Cues: Implicit Video Question Answering via Dual-Clue Reasoning
von: Chen, Tieyuan, et al.
Veröffentlicht: (2025)
von: Chen, Tieyuan, et al.
Veröffentlicht: (2025)
Unleashing Diffusion Transformers for Visual Correspondence by Modulating Massive Activations
von: Gan, Chaofan, et al.
Veröffentlicht: (2025)
von: Gan, Chaofan, et al.
Veröffentlicht: (2025)
MECD+: Unlocking Event-Level Causal Graph Discovery for Video Reasoning
von: Chen, Tieyuan, et al.
Veröffentlicht: (2025)
von: Chen, Tieyuan, et al.
Veröffentlicht: (2025)
ImageFolder: Autoregressive Image Generation with Folded Tokens
von: Li, Xiang, et al.
Veröffentlicht: (2024)
von: Li, Xiang, et al.
Veröffentlicht: (2024)
Enhancing Video Large Language Models with Structured Multi-Video Collaborative Reasoning
von: He, Zhihao, et al.
Veröffentlicht: (2025)
von: He, Zhihao, et al.
Veröffentlicht: (2025)
AccDiffusion: An Accurate Method for Higher-Resolution Image Generation
von: Lin, Zhihang, et al.
Veröffentlicht: (2024)
von: Lin, Zhihang, et al.
Veröffentlicht: (2024)
Parallel Vision Token Scheduling for Fast and Accurate Multimodal LMMs Inference
von: Zhan, Wengyi, et al.
Veröffentlicht: (2025)
von: Zhan, Wengyi, et al.
Veröffentlicht: (2025)
Boosting Multimodal Large Language Models with Visual Tokens Withdrawal for Rapid Inference
von: Lin, Zhihang, et al.
Veröffentlicht: (2024)
von: Lin, Zhihang, et al.
Veröffentlicht: (2024)
MECD: Unlocking Multi-Event Causal Discovery in Video Reasoning
von: Chen, Tieyuan, et al.
Veröffentlicht: (2024)
von: Chen, Tieyuan, et al.
Veröffentlicht: (2024)
X-Cache: Cross-Chunk Block Caching for Few-Step Autoregressive World Models Inference
von: Zeng, Yixiao, et al.
Veröffentlicht: (2026)
von: Zeng, Yixiao, et al.
Veröffentlicht: (2026)
XQ-GAN: An Open-source Image Tokenization Framework for Autoregressive Generation
von: Li, Xiang, et al.
Veröffentlicht: (2024)
von: Li, Xiang, et al.
Veröffentlicht: (2024)
CSTA: Spatial-Temporal Causal Adaptive Learning for Exemplar-Free Video Class-Incremental Learning
von: Chen, Tieyuan, et al.
Veröffentlicht: (2025)
von: Chen, Tieyuan, et al.
Veröffentlicht: (2025)
Move and Act: Enhanced Object Manipulation and Background Integrity for Image Editing
von: Jiang, Pengfei, et al.
Veröffentlicht: (2024)
von: Jiang, Pengfei, et al.
Veröffentlicht: (2024)
Test-Time Temporal Sampling for Efficient MLLM Video Understanding
von: Wang, Kaibin, et al.
Veröffentlicht: (2025)
von: Wang, Kaibin, et al.
Veröffentlicht: (2025)
Progressive Supernet Training for Efficient Visual Autoregressive Modeling
von: Chen, Xiaoyue, et al.
Veröffentlicht: (2025)
von: Chen, Xiaoyue, et al.
Veröffentlicht: (2025)
VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation
von: Lin, Huawei, et al.
Veröffentlicht: (2025)
von: Lin, Huawei, et al.
Veröffentlicht: (2025)
Improving Flexible Image Tokenizers for Autoregressive Image Generation
von: Fu, Zixuan, et al.
Veröffentlicht: (2026)
von: Fu, Zixuan, et al.
Veröffentlicht: (2026)
ALTo: Adaptive-Length Tokenizer for Autoregressive Mask Generation
von: Wang, Lingfeng, et al.
Veröffentlicht: (2025)
von: Wang, Lingfeng, et al.
Veröffentlicht: (2025)
Image Tokenizer Needs Post-Training
von: Qiu, Kai, et al.
Veröffentlicht: (2025)
von: Qiu, Kai, et al.
Veröffentlicht: (2025)
FastVAR: Linear Visual Autoregressive Modeling via Cached Token Pruning
von: Guo, Hang, et al.
Veröffentlicht: (2025)
von: Guo, Hang, et al.
Veröffentlicht: (2025)
VibeToken: Scaling 1D Image Tokenizers and Autoregressive Models for Dynamic Resolution Generations
von: Patel, Maitreya, et al.
Veröffentlicht: (2026)
von: Patel, Maitreya, et al.
Veröffentlicht: (2026)
From Priors to Perception: Grounding Video-LLMs in Physical Reality
von: Zhao, Zicheng, et al.
Veröffentlicht: (2026)
von: Zhao, Zicheng, et al.
Veröffentlicht: (2026)
DAC: 2D-3D Retrieval with Noisy Labels via Divide-and-Conquer Alignment and Correction
von: Gan, Chaofan, et al.
Veröffentlicht: (2024)
von: Gan, Chaofan, et al.
Veröffentlicht: (2024)
Bridging Continuous and Discrete Tokens for Autoregressive Visual Generation
von: Wang, Yuqing, et al.
Veröffentlicht: (2025)
von: Wang, Yuqing, et al.
Veröffentlicht: (2025)
Hita: Holistic Tokenizer for Autoregressive Image Generation
von: Zheng, Anlin, et al.
Veröffentlicht: (2025)
von: Zheng, Anlin, et al.
Veröffentlicht: (2025)
NextStep-1: Toward Autoregressive Image Generation with Continuous Tokens at Scale
von: NextStep Team, et al.
Veröffentlicht: (2025)
von: NextStep Team, et al.
Veröffentlicht: (2025)
Improving Autoregressive Image Generation through Coarse-to-Fine Token Prediction
von: Guo, Ziyao, et al.
Veröffentlicht: (2025)
von: Guo, Ziyao, et al.
Veröffentlicht: (2025)
ObjectAdd: Adding Objects into Image via a Training-Free Diffusion Modification Fashion
von: Zhang, Ziyue, et al.
Veröffentlicht: (2024)
von: Zhang, Ziyue, et al.
Veröffentlicht: (2024)
Compute Only 16 Tokens in One Timestep: Accelerating Diffusion Transformers with Cluster-Driven Feature Caching
von: Zheng, Zhixin, et al.
Veröffentlicht: (2025)
von: Zheng, Zhixin, et al.
Veröffentlicht: (2025)
LINA: Linear Autoregressive Image Generative Models with Continuous Tokens
von: Wang, Jiahao, et al.
Veröffentlicht: (2026)
von: Wang, Jiahao, et al.
Veröffentlicht: (2026)
You Only Need Less Attention at Each Stage in Vision Transformers
von: Zhang, Shuoxi, et al.
Veröffentlicht: (2024)
von: Zhang, Shuoxi, et al.
Veröffentlicht: (2024)
EVATok: Adaptive Length Video Tokenization for Efficient Visual Autoregressive Generation
von: Xiong, Tianwei, et al.
Veröffentlicht: (2026)
von: Xiong, Tianwei, et al.
Veröffentlicht: (2026)
Frequency Autoregressive Image Generation with Continuous Tokens
von: Yu, Hu, et al.
Veröffentlicht: (2025)
von: Yu, Hu, et al.
Veröffentlicht: (2025)
AnySR: Realizing Image Super-Resolution as Any-Scale, Any-Resource
von: Zhan, Wengyi, et al.
Veröffentlicht: (2024)
von: Zhan, Wengyi, et al.
Veröffentlicht: (2024)
Token-Shuffle: Towards High-Resolution Image Generation with Autoregressive Models
von: Ma, Xu, et al.
Veröffentlicht: (2025)
von: Ma, Xu, et al.
Veröffentlicht: (2025)
Treat Visual Tokens as Text? But Your MLLM Only Needs Fewer Efforts to See
von: Zhang, Zeliang, et al.
Veröffentlicht: (2024)
von: Zhang, Zeliang, et al.
Veröffentlicht: (2024)
EasyInv: Toward Fast and Better DDIM Inversion
von: Zhang, Ziyue, et al.
Veröffentlicht: (2024)
von: Zhang, Ziyue, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Head-Aware KV Cache Compression for Efficient Visual Autoregressive Modeling
von: Qin, Ziran, et al.
Veröffentlicht: (2025) -
Massive Activations are the Key to Local Detail Synthesis in Diffusion Transformers
von: Gan, Chaofan, et al.
Veröffentlicht: (2025) -
VidLaDA: Bidirectional Diffusion Large Language Models for Efficient Video Understanding
von: He, Zhihao, et al.
Veröffentlicht: (2026) -
Looking Beyond Visible Cues: Implicit Video Question Answering via Dual-Clue Reasoning
von: Chen, Tieyuan, et al.
Veröffentlicht: (2025) -
Unleashing Diffusion Transformers for Visual Correspondence by Modulating Massive Activations
von: Gan, Chaofan, et al.
Veröffentlicht: (2025)