Visual Autoregressive Transformers Must Use $Ω(n^2 d)$ Memory
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Cao, Yang, Li, Xiaoyu, Ke, Yekun, Liang, Yingyu, Shi, Zhenmei, Song, Zhao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
On Computational Limits and Provably Efficient Criteria of Visual Autoregressive Models: A Fine-Grained Complexity Analysis
von: Ke, Yekun, et al.
Veröffentlicht: (2025)
von: Ke, Yekun, et al.
Veröffentlicht: (2025)
Universal Approximation of Visual Autoregressive Transformers
von: Chen, Yifang, et al.
Veröffentlicht: (2025)
von: Chen, Yifang, et al.
Veröffentlicht: (2025)
HOFAR: High-Order Augmentation of Flow Autoregressive Transformers
von: Liang, Yingyu, et al.
Veröffentlicht: (2025)
von: Liang, Yingyu, et al.
Veröffentlicht: (2025)
On Computational Limits of FlowAR Models: Expressivity and Efficiency
von: Cao, Yang, et al.
Veröffentlicht: (2025)
von: Cao, Yang, et al.
Veröffentlicht: (2025)
Circuit Complexity Bounds for Visual Autoregressive Model
von: Ke, Yekun, et al.
Veröffentlicht: (2025)
von: Ke, Yekun, et al.
Veröffentlicht: (2025)
Unraveling the Smoothness Properties of Diffusion Models: A Gaussian Mixture Perspective
von: Liang, Yingyu, et al.
Veröffentlicht: (2024)
von: Liang, Yingyu, et al.
Veröffentlicht: (2024)
Force Matching with Relativistic Constraints: A Physics-Inspired Approach to Stable and Efficient Generative Modeling
von: Cao, Yang, et al.
Veröffentlicht: (2025)
von: Cao, Yang, et al.
Veröffentlicht: (2025)
RichSpace: Enriching Text-to-Video Prompt Space via Text Embedding Interpolation
von: Cao, Yuefan, et al.
Veröffentlicht: (2025)
von: Cao, Yuefan, et al.
Veröffentlicht: (2025)
Theoretical Guarantees for High Order Trajectory Refinement in Generative Flows
von: Gong, Chengyue, et al.
Veröffentlicht: (2025)
von: Gong, Chengyue, et al.
Veröffentlicht: (2025)
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond
von: Ke, Yekun, et al.
Veröffentlicht: (2024)
von: Ke, Yekun, et al.
Veröffentlicht: (2024)
High-Order Matching for One-Step Shortcut Diffusion Models
von: Chen, Bo, et al.
Veröffentlicht: (2025)
von: Chen, Bo, et al.
Veröffentlicht: (2025)
Advancing the Understanding of Fixed Point Iterations in Deep Neural Networks: A Detailed Analytical Study
von: Ke, Yekun, et al.
Veröffentlicht: (2024)
von: Ke, Yekun, et al.
Veröffentlicht: (2024)
Text-to-Image Diffusion Models Cannot Count, and Prompt Refinement Cannot Help
von: Guo, Xuyang, et al.
Veröffentlicht: (2025)
von: Guo, Xuyang, et al.
Veröffentlicht: (2025)
Can You Count to Nine? A Human Evaluation Benchmark for Counting Limits in Modern Text-to-Video Models
von: Guo, Xuyang, et al.
Veröffentlicht: (2025)
von: Guo, Xuyang, et al.
Veröffentlicht: (2025)
Provable Differentially Private Computation of the Cross-Attention Mechanism
von: Ke, Yekun, et al.
Veröffentlicht: (2024)
von: Ke, Yekun, et al.
Veröffentlicht: (2024)
Your Vision-Language Model Can't Even Count to 20: Exposing the Failures of VLMs in Compositional Counting
von: Guo, Xuyang, et al.
Veröffentlicht: (2025)
von: Guo, Xuyang, et al.
Veröffentlicht: (2025)
HART: Efficient Visual Generation with Hybrid Autoregressive Transformer
von: Tang, Haotian, et al.
Veröffentlicht: (2024)
von: Tang, Haotian, et al.
Veröffentlicht: (2024)
Spanning Tree Autoregressive Visual Generation
von: Lee, Sangkyu, et al.
Veröffentlicht: (2025)
von: Lee, Sangkyu, et al.
Veröffentlicht: (2025)
TokenUnify: Scaling Up Autoregressive Pretraining for Neuron Segmentation
von: Chen, Yinda, et al.
Veröffentlicht: (2024)
von: Chen, Yinda, et al.
Veröffentlicht: (2024)
Trustworthy Hate Speech Detection Through Visual Augmentation
von: Yang, Ziyuan, et al.
Veröffentlicht: (2024)
von: Yang, Ziyuan, et al.
Veröffentlicht: (2024)
Revisiting Non-Autoregressive Transformers for Efficient Image Synthesis
von: Ni, Zanlin, et al.
Veröffentlicht: (2024)
von: Ni, Zanlin, et al.
Veröffentlicht: (2024)
LANTERN: Accelerating Visual Autoregressive Models with Relaxed Speculative Decoding
von: Jang, Doohyuk, et al.
Veröffentlicht: (2024)
von: Jang, Doohyuk, et al.
Veröffentlicht: (2024)
Adversarial Error Correction for Visual Autoregressive Generation
von: Bi, Ligong, et al.
Veröffentlicht: (2026)
von: Bi, Ligong, et al.
Veröffentlicht: (2026)
Generative Pre-trained Autoregressive Diffusion Transformer
von: Zhang, Yuan, et al.
Veröffentlicht: (2025)
von: Zhang, Yuan, et al.
Veröffentlicht: (2025)
MeshAnything: Artist-Created Mesh Generation with Autoregressive Transformers
von: Chen, Yiwen, et al.
Veröffentlicht: (2024)
von: Chen, Yiwen, et al.
Veröffentlicht: (2024)
ControlVAR: Exploring Controllable Visual Autoregressive Modeling
von: Li, Xiang, et al.
Veröffentlicht: (2024)
von: Li, Xiang, et al.
Veröffentlicht: (2024)
StyleVAR: Controllable Image Style Transfer via Visual Autoregressive Modeling
von: Jing, Liqi, et al.
Veröffentlicht: (2026)
von: Jing, Liqi, et al.
Veröffentlicht: (2026)
Towards Infinite-Long Prefix in Transformer
von: Liang, Yingyu, et al.
Veröffentlicht: (2024)
von: Liang, Yingyu, et al.
Veröffentlicht: (2024)
Bypassing the Exponential Dependency: Looped Transformers Efficiently Learn In-context by Multi-step Gradient Descent
von: Chen, Bo, et al.
Veröffentlicht: (2024)
von: Chen, Bo, et al.
Veröffentlicht: (2024)
Exploring the Frontiers of Softmax: Provable Optimization, Applications in Diffusion Model, and Beyond
von: Cao, Yang, et al.
Veröffentlicht: (2024)
von: Cao, Yang, et al.
Veröffentlicht: (2024)
VARGPT-v1.1: Improve Visual Autoregressive Large Unified Model via Iterative Instruction Tuning and Reinforcement Learning
von: Zhuang, Xianwei, et al.
Veröffentlicht: (2025)
von: Zhuang, Xianwei, et al.
Veröffentlicht: (2025)
Beyond Visual Memory: Mechanistic Diagnostics of Latent Visual Reasoning
von: Guo, Garvin, et al.
Veröffentlicht: (2026)
von: Guo, Garvin, et al.
Veröffentlicht: (2026)
Time and Memory Trade-off of KV-Cache Compression in Tensor Transformer Decoding
von: Chen, Yifang, et al.
Veröffentlicht: (2025)
von: Chen, Yifang, et al.
Veröffentlicht: (2025)
TransNeXt: Robust Foveal Visual Perception for Vision Transformers
von: Shi, Dai
Veröffentlicht: (2023)
von: Shi, Dai
Veröffentlicht: (2023)
CSD-VAR: Content-Style Decomposition in Visual Autoregressive Models
von: Nguyen, Quang-Binh, et al.
Veröffentlicht: (2025)
von: Nguyen, Quang-Binh, et al.
Veröffentlicht: (2025)
MAGI-1: Autoregressive Video Generation at Scale
von: ai, Sand., et al.
Veröffentlicht: (2025)
von: ai, Sand., et al.
Veröffentlicht: (2025)
VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation
von: Lin, Huawei, et al.
Veröffentlicht: (2025)
von: Lin, Huawei, et al.
Veröffentlicht: (2025)
T2VTextBench: A Human Evaluation Benchmark for Textual Control in Video Generation Models
von: Guo, Xuyang, et al.
Veröffentlicht: (2025)
von: Guo, Xuyang, et al.
Veröffentlicht: (2025)
T2VPhysBench: A First-Principles Benchmark for Physical Consistency in Text-to-Video Generation
von: Guo, Xuyang, et al.
Veröffentlicht: (2025)
von: Guo, Xuyang, et al.
Veröffentlicht: (2025)
Vision Foundation Models as Effective Visual Tokenizers for Autoregressive Image Generation
von: Zheng, Anlin, et al.
Veröffentlicht: (2025)
von: Zheng, Anlin, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
On Computational Limits and Provably Efficient Criteria of Visual Autoregressive Models: A Fine-Grained Complexity Analysis
von: Ke, Yekun, et al.
Veröffentlicht: (2025) -
Universal Approximation of Visual Autoregressive Transformers
von: Chen, Yifang, et al.
Veröffentlicht: (2025) -
HOFAR: High-Order Augmentation of Flow Autoregressive Transformers
von: Liang, Yingyu, et al.
Veröffentlicht: (2025) -
On Computational Limits of FlowAR Models: Expressivity and Efficiency
von: Cao, Yang, et al.
Veröffentlicht: (2025) -
Circuit Complexity Bounds for Visual Autoregressive Model
von: Ke, Yekun, et al.
Veröffentlicht: (2025)