Token-Shuffle: Towards High-Resolution Image Generation with Autoregressive Models
Fuente:
arXiv
Saved in:
| Main Authors: | Ma, Xu, Sun, Peize, Ma, Haoyu, Tang, Hao, Ma, Chih-Yao, Wang, Jialiang, Li, Kunpeng, Dai, Xiaoliang, Shi, Yujun, Ju, Xuan, Hu, Yushi, Sanakoyeu, Artsiom, Juefei-Xu, Felix, Hou, Ji, Tian, Junjiao, Xu, Tao, Hou, Tingbo, Liu, Yen-Cheng, He, Zecheng, He, Zijian, Feiszli, Matt, Zhang, Peizhao, Vajda, Peter, Tsai, Sam, Fu, Yun |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Movie Weaver: Tuning-Free Multi-Concept Video Personalization with Anchored Prompts
by: Liang, Feng, et al.
Published: (2025)
by: Liang, Feng, et al.
Published: (2025)
Populate-A-Scene: Affordance-Aware Human Video Generation
by: Shan, Mengyi, et al.
Published: (2025)
by: Shan, Mengyi, et al.
Published: (2025)
Llama Learns to Direct: DirectorLLM for Human-Centric Video Generation
by: Song, Kunpeng, et al.
Published: (2024)
by: Song, Kunpeng, et al.
Published: (2024)
Non-Markov Multi-Round Conversational Image Generation with History-Conditioned MLLMs
by: Zhang, Haochen, et al.
Published: (2026)
by: Zhang, Haochen, et al.
Published: (2026)
MoCha: Towards Movie-Grade Talking Character Synthesis
by: Wei, Cong, et al.
Published: (2025)
by: Wei, Cong, et al.
Published: (2025)
Cache Me if You Can: Accelerating Diffusion Models through Block Caching
by: Wimbauer, Felix, et al.
Published: (2023)
by: Wimbauer, Felix, et al.
Published: (2023)
Improving Chain-of-Thought Efficiency for Autoregressive Image Generation
by: Gu, Zeqi, et al.
Published: (2025)
by: Gu, Zeqi, et al.
Published: (2025)
LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity
by: Wang, Hongjie, et al.
Published: (2024)
by: Wang, Hongjie, et al.
Published: (2024)
Pixel-Space Post-Training of Latent Diffusion Models
by: Zhang, Christina, et al.
Published: (2024)
by: Zhang, Christina, et al.
Published: (2024)
Imagine Flash: Accelerating Emu Diffusion Models with Backward Distillation
by: Kohler, Jonas, et al.
Published: (2024)
by: Kohler, Jonas, et al.
Published: (2024)
Think in Strokes, Not Pixels: Process-Driven Image Generation via Interleaved Reasoning
by: Zhang, Lei, et al.
Published: (2026)
by: Zhang, Lei, et al.
Published: (2026)
StreamDiT: Real-Time Streaming Text-to-Video Generation
by: Kodaira, Akio, et al.
Published: (2025)
by: Kodaira, Akio, et al.
Published: (2025)
Imagine yourself: Tuning-Free Personalized Image Generation
by: He, Zecheng, et al.
Published: (2024)
by: He, Zecheng, et al.
Published: (2024)
Autoregressive Distillation of Diffusion Transformers
by: Kim, Yeongmin, et al.
Published: (2025)
by: Kim, Yeongmin, et al.
Published: (2025)
UniT: Unified Multimodal Chain-of-Thought Test-time Scaling
by: Chen, Leon Liangyu, et al.
Published: (2026)
by: Chen, Leon Liangyu, et al.
Published: (2026)
Learnings from Scaling Visual Tokenizers for Reconstruction and Generation
by: Hansen-Estruch, Philippe, et al.
Published: (2025)
by: Hansen-Estruch, Philippe, et al.
Published: (2025)
FBNetV5: Neural Architecture Search for Multiple Tasks in One Run
by: Wu, Bichen, et al.
Published: (2021)
by: Wu, Bichen, et al.
Published: (2021)
An Analysis on Quantizing Diffusion Transformers
by: Yang, Yuewei, et al.
Published: (2024)
by: Yang, Yuewei, et al.
Published: (2024)
Using Motion Cues to Supervise Single-Frame Body Pose and Shape Estimation in Low Data Regimes
by: Davydov, Andrey, et al.
Published: (2024)
by: Davydov, Andrey, et al.
Published: (2024)
On the Bivariate Characteristic Polynomial of the Shuffle Lattice
by: Ma, Annabel
Published: (2024)
by: Ma, Annabel
Published: (2024)
CESM coupled simulation dataset with new (Mod_cp) and original (F09) dynamical coupling scheme between atmosphere and ocean
by: Ma, Jialiang, et al.
Published: (2020)
by: Ma, Jialiang, et al.
Published: (2020)
FlowVid: Taming Imperfect Optical Flows for Consistent Video-to-Video Synthesis
by: Liang, Feng, et al.
Published: (2023)
by: Liang, Feng, et al.
Published: (2023)
Exploring MLLM-Diffusion Information Transfer with MetaCanvas
by: Lin, Han, et al.
Published: (2025)
by: Lin, Han, et al.
Published: (2025)
PhyGDPO: Physics-Aware Groupwise Direct Preference Optimization for Physically Consistent Text-to-Video Generation
by: Cai, Yuanhao, et al.
Published: (2025)
by: Cai, Yuanhao, et al.
Published: (2025)
MobileDiffusion: Instant Text-to-Image Generation on Mobile Devices
by: Zhao, Yang, et al.
Published: (2023)
by: Zhao, Yang, et al.
Published: (2023)
Castling-ViT: Compressing Self-Attention via Switching Towards Linear-Angular Attention at Vision Transformer Inference
by: You, Haoran, et al.
Published: (2022)
by: You, Haoran, et al.
Published: (2022)
FLIP: Real-Time and Resilient Formation Planning for Large-Scale DIstributed Swarms via Point Cloud Registration
by: Zhou, Yuan, et al.
Published: (2026)
by: Zhou, Yuan, et al.
Published: (2026)
Transfer between Modalities with MetaQueries
by: Pan, Xichen, et al.
Published: (2025)
by: Pan, Xichen, et al.
Published: (2025)
Message Passing Variational Autoregressive Network for Solving Intractable Ising Models
by: Ma, Qunlong, et al.
Published: (2024)
by: Ma, Qunlong, et al.
Published: (2024)
Number Adaptive Formation Flight Planning via Affine Deformable Guidance in Narrow Environments
by: Zhou, Yuan, et al.
Published: (2025)
by: Zhou, Yuan, et al.
Published: (2025)
Eliciting Chain-of-Thought in Base LLMs via Gradient-Based Representation Optimization
by: Wang, Zijian, et al.
Published: (2025)
by: Wang, Zijian, et al.
Published: (2025)
Delay-Doppler Domain Channel Estimation: What if Sparsity is Unknown?
by: Yang, Zijian, et al.
Published: (2026)
by: Yang, Zijian, et al.
Published: (2026)
ADen: Adaptive Density Representations for Sparse-view Camera Pose Estimation
by: Tang, Hao, et al.
Published: (2024)
by: Tang, Hao, et al.
Published: (2024)
AcTracer: Active Testing of Large Language Model via Multi-Stage Sampling
by: Huang, Yuheng, et al.
Published: (2024)
by: Huang, Yuheng, et al.
Published: (2024)
Unleashing In-context Learning of Autoregressive Models for Few-shot Image Manipulation
by: Lai, Bolin, et al.
Published: (2024)
by: Lai, Bolin, et al.
Published: (2024)
Multi-relational Network Autoregression Model with Latent Group Structures
by: Ren, Yimeng, et al.
Published: (2024)
by: Ren, Yimeng, et al.
Published: (2024)
INGeo: Accelerating Instant Neural Scene Reconstruction with Noisy Geometry Priors
by: Li, Chaojian, et al.
Published: (2022)
by: Li, Chaojian, et al.
Published: (2022)
Next Block Prediction: Video Generation via Semi-Autoregressive Modeling
by: Ren, Shuhuai, et al.
Published: (2025)
by: Ren, Shuhuai, et al.
Published: (2025)
On the Last-Iterate Convergence of Shuffling Gradient Methods
by: Liu, Zijian, et al.
Published: (2024)
by: Liu, Zijian, et al.
Published: (2024)
Counting triangles in regular graphs
by: He, Jialin, et al.
Published: (2023)
by: He, Jialin, et al.
Published: (2023)
Similar Items
-
Movie Weaver: Tuning-Free Multi-Concept Video Personalization with Anchored Prompts
by: Liang, Feng, et al.
Published: (2025) -
Populate-A-Scene: Affordance-Aware Human Video Generation
by: Shan, Mengyi, et al.
Published: (2025) -
Llama Learns to Direct: DirectorLLM for Human-Centric Video Generation
by: Song, Kunpeng, et al.
Published: (2024) -
Non-Markov Multi-Round Conversational Image Generation with History-Conditioned MLLMs
by: Zhang, Haochen, et al.
Published: (2026) -
MoCha: Towards Movie-Grade Talking Character Synthesis
by: Wei, Cong, et al.
Published: (2025)