Understand Before You Generate: Self-Guided Training for Autoregressive Image Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Yue, Xiaoyu, Wang, Zidong, Wang, Yuqing, Zhang, Wenlong, Liu, Xihui, Ouyang, Wanli, Bai, Lei, Zhou, Luping |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Diffusion Models Need Visual Priors for Image Generation
by: Yue, Xiaoyu, et al.
Published: (2024)
by: Yue, Xiaoyu, et al.
Published: (2024)
Transition Models: Rethinking the Generative Learning Objective
by: Wang, Zidong, et al.
Published: (2025)
by: Wang, Zidong, et al.
Published: (2025)
Native-Resolution Image Synthesis
by: Wang, Zidong, et al.
Published: (2025)
by: Wang, Zidong, et al.
Published: (2025)
FiT: Flexible Vision Transformer for Diffusion Model
by: Lu, Zeyu, et al.
Published: (2024)
by: Lu, Zeyu, et al.
Published: (2024)
Bridging Continuous and Discrete Tokens for Autoregressive Visual Generation
by: Wang, Yuqing, et al.
Published: (2025)
by: Wang, Yuqing, et al.
Published: (2025)
Parallelized Autoregressive Visual Generation
by: Wang, Yuqing, et al.
Published: (2024)
by: Wang, Yuqing, et al.
Published: (2024)
Loong: Generating Minute-level Long Videos with Autoregressive Language Models
by: Wang, Yuqing, et al.
Published: (2024)
by: Wang, Yuqing, et al.
Published: (2024)
MACRO: Advancing Multi-Reference Image Generation with Structured Long-Context Data
by: Chen, Zhekai, et al.
Published: (2026)
by: Chen, Zhekai, et al.
Published: (2026)
A Renaissance of Explicit Motion Information Mining from Transformers for Action Recognition
by: Zhuang, Peiqin, et al.
Published: (2025)
by: Zhuang, Peiqin, et al.
Published: (2025)
Improving Weakly Supervised Temporal Action Localization by Exploiting Multi-resolution Information in Temporal Domain
by: Su, Rui, et al.
Published: (2025)
by: Su, Rui, et al.
Published: (2025)
Progressive Cross-Stream Cooperation in Spatial and Temporal Domain for Action Localization
by: Su, Rui, et al.
Published: (2019)
by: Su, Rui, et al.
Published: (2019)
AR-GRPO: Training Autoregressive Image Generation Models via Reinforcement Learning
by: Yuan, Shihao, et al.
Published: (2025)
by: Yuan, Shihao, et al.
Published: (2025)
InfGen: A Resolution-Agnostic Paradigm for Scalable Image Synthesis
by: Han, Tao, et al.
Published: (2025)
by: Han, Tao, et al.
Published: (2025)
FlashAR: Efficient Post-Training Acceleration for Autoregressive Image Generation
by: Zhou, Junkang, et al.
Published: (2026)
by: Zhou, Junkang, et al.
Published: (2026)
Dense Video Captioning using Graph-based Sentence Summarization
by: Zhang, Zhiwang, et al.
Published: (2025)
by: Zhang, Zhiwang, et al.
Published: (2025)
ComfyBench: Benchmarking LLM-based Agents in ComfyUI for Autonomously Designing Collaborative AI Systems
by: Xue, Xiangyuan, et al.
Published: (2024)
by: Xue, Xiangyuan, et al.
Published: (2024)
GigaTok: Scaling Visual Tokenizers to 3 Billion Parameters for Autoregressive Image Generation
by: Xiong, Tianwei, et al.
Published: (2025)
by: Xiong, Tianwei, et al.
Published: (2025)
Exploring Representation-Aligned Latent Space for Better Generation
by: Xu, Wanghan, et al.
Published: (2025)
by: Xu, Wanghan, et al.
Published: (2025)
WorldSimBench: Towards Video Generation Models as World Simulators
by: Qin, Yiran, et al.
Published: (2024)
by: Qin, Yiran, et al.
Published: (2024)
Conditional Panoramic Image Generation via Masked Autoregressive Modeling
by: Wang, Chaoyang, et al.
Published: (2025)
by: Wang, Chaoyang, et al.
Published: (2025)
PredBench: Benchmarking Spatio-Temporal Prediction across Diverse Disciplines
by: Wang, ZiDong, et al.
Published: (2024)
by: Wang, ZiDong, et al.
Published: (2024)
RoboMirror: Understand Before You Imitate for Video to Humanoid Locomotion
by: Li, Zhe, et al.
Published: (2025)
by: Li, Zhe, et al.
Published: (2025)
Divide and Conquer: Language Models can Plan and Self-Correct for Compositional Text-to-Image Generation
by: Wang, Zhenyu, et al.
Published: (2024)
by: Wang, Zhenyu, et al.
Published: (2024)
Think Before You Move: Latent Motion Reasoning for Text-to-Motion Generation
by: Qian, Yijie, et al.
Published: (2025)
by: Qian, Yijie, et al.
Published: (2025)
SATURN: Autoregressive Image Generation Guided by Scene Graphs
by: Vo, Thanh-Nhan, et al.
Published: (2025)
by: Vo, Thanh-Nhan, et al.
Published: (2025)
MedXChat: A Unified Multimodal Large Language Model Framework towards CXRs Understanding and Generation
by: Yang, Ling, et al.
Published: (2023)
by: Yang, Ling, et al.
Published: (2023)
MotionGPT-2: A General-Purpose Motion-Language Model for Motion Generation and Understanding
by: Wang, Yuan, et al.
Published: (2024)
by: Wang, Yuan, et al.
Published: (2024)
ReconMOST: Multi-Layer Sea Temperature Reconstruction with Observations-Guided Diffusion
by: Song, Yuanyi, et al.
Published: (2025)
by: Song, Yuanyi, et al.
Published: (2025)
Speculative Jacobi-Denoising Decoding for Accelerating Autoregressive Text-to-image Generation
by: Teng, Yao, et al.
Published: (2025)
by: Teng, Yao, et al.
Published: (2025)
Perceive, Understand and Restore: Real-World Image Super-Resolution with Autoregressive Multimodal Generative Models
by: Wei, Hongyang, et al.
Published: (2025)
by: Wei, Hongyang, et al.
Published: (2025)
Training-Free Watermarking for Autoregressive Image Generation
by: Tong, Yu, et al.
Published: (2025)
by: Tong, Yu, et al.
Published: (2025)
ARGenSeg: Image Segmentation with Autoregressive Image Generation Model
by: Wang, Xiaolong, et al.
Published: (2025)
by: Wang, Xiaolong, et al.
Published: (2025)
GenArtist: Multimodal LLM as an Agent for Unified Image Generation and Editing
by: Wang, Zhenyu, et al.
Published: (2024)
by: Wang, Zhenyu, et al.
Published: (2024)
EVATok: Adaptive Length Video Tokenization for Efficient Visual Autoregressive Generation
by: Xiong, Tianwei, et al.
Published: (2026)
by: Xiong, Tianwei, et al.
Published: (2026)
UniX: Unifying Autoregression and Diffusion for Chest X-Ray Understanding and Generation
by: Zhang, Ruiheng, et al.
Published: (2026)
by: Zhang, Ruiheng, et al.
Published: (2026)
Training-Free Text-Guided Image Editing with Visual Autoregressive Model
by: Wang, Yufei, et al.
Published: (2025)
by: Wang, Yufei, et al.
Published: (2025)
Scalable Autoregressive Image Generation with Mamba
by: Li, Haopeng, et al.
Published: (2024)
by: Li, Haopeng, et al.
Published: (2024)
RS-Mamba for Large Remote Sensing Image Dense Prediction
by: Zhao, Sijie, et al.
Published: (2024)
by: Zhao, Sijie, et al.
Published: (2024)
Autoregressive Omni-Aware Outpainting for Open-Vocabulary 360-Degree Image Generation
by: Lu, Zhuqiang, et al.
Published: (2023)
by: Lu, Zhuqiang, et al.
Published: (2023)
Autoregressive Image Generation with Randomized Parallel Decoding
by: Li, Haopeng, et al.
Published: (2025)
by: Li, Haopeng, et al.
Published: (2025)
Similar Items
-
Diffusion Models Need Visual Priors for Image Generation
by: Yue, Xiaoyu, et al.
Published: (2024) -
Transition Models: Rethinking the Generative Learning Objective
by: Wang, Zidong, et al.
Published: (2025) -
Native-Resolution Image Synthesis
by: Wang, Zidong, et al.
Published: (2025) -
FiT: Flexible Vision Transformer for Diffusion Model
by: Lu, Zeyu, et al.
Published: (2024) -
Bridging Continuous and Discrete Tokens for Autoregressive Visual Generation
by: Wang, Yuqing, et al.
Published: (2025)