VPG: Visual Prefix Guidance for Autoregressive Image and Video Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Liao, Xinyao, He, Qiyuan, Li, Yicong, Zhu, Jiayin, Qu, Xiaoye, Wei, Wei, Yao, Angela |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VA-$π$: Variational Policy Alignment for Pixel-Aware Autoregressive Generation
by: Liao, Xinyao, et al.
Published: (2025)
by: Liao, Xinyao, et al.
Published: (2025)
REAR: Rethinking Visual Autoregressive Models via Generator-Tokenizer Consistency Regularization
by: He, Qiyuan, et al.
Published: (2025)
by: He, Qiyuan, et al.
Published: (2025)
Interp3D: Correspondence-aware Interpolation for Generative Textured 3D Morphing
by: Liu, Xiaolu, et al.
Published: (2026)
by: Liu, Xiaolu, et al.
Published: (2026)
RelaxFlow: Text-Driven Amodal 3D Generation
by: Zhu, Jiayin, et al.
Published: (2026)
by: Zhu, Jiayin, et al.
Published: (2026)
Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning
by: Liao, Xinyao, et al.
Published: (2025)
by: Liao, Xinyao, et al.
Published: (2025)
AnchorDS: Anchoring Dynamic Sources for Semantically Consistent Text-to-3D Generation
by: Zhu, Jiayin, et al.
Published: (2025)
by: Zhu, Jiayin, et al.
Published: (2025)
Conceptrol: Concept Control of Zero-shot Personalized Image Generation
by: He, Qiyuan, et al.
Published: (2025)
by: He, Qiyuan, et al.
Published: (2025)
Does Engram Do Memory Retrieval in Autoregressive Image Generation?
by: Wang, Jinghao, et al.
Published: (2026)
by: Wang, Jinghao, et al.
Published: (2026)
Alleviating Hallucination in Large Vision-Language Models with Active Retrieval Augmentation
by: Qu, Xiaoye, et al.
Published: (2024)
by: Qu, Xiaoye, et al.
Published: (2024)
Can I Trust Your Answer? Visually Grounded Video Question Answering
by: Xiao, Junbin, et al.
Published: (2023)
by: Xiao, Junbin, et al.
Published: (2023)
SATORI-R1: Incentivizing Multimodal Reasoning through Explicit Visual Anchoring
by: Shen, Chuming, et al.
Published: (2025)
by: Shen, Chuming, et al.
Published: (2025)
InterDyad: Interactive Dyadic Speech-to-Video Generation by Querying Intermediate Visual Guidance
by: Pan, Dongwei, et al.
Published: (2026)
by: Pan, Dongwei, et al.
Published: (2026)
Pushing the Boundaries of State Space Models for Image and Video Generation
by: Hong, Yicong, et al.
Published: (2025)
by: Hong, Yicong, et al.
Published: (2025)
SSG: Scaled Spatial Guidance for Multi-Scale Visual Autoregressive Generation
by: Shin, Youngwoo, et al.
Published: (2026)
by: Shin, Youngwoo, et al.
Published: (2026)
Semantic-Aware Prefix Learning for Token-Efficient Image Generation
by: Li, Qingfeng, et al.
Published: (2026)
by: Li, Qingfeng, et al.
Published: (2026)
VCE: Safe Autoregressive Image Generation via Visual Contrast Exploitation
by: Han, Feng, et al.
Published: (2025)
by: Han, Feng, et al.
Published: (2025)
VideoSSR: Video Self-Supervised Reinforcement Learning
by: He, Zefeng, et al.
Published: (2025)
by: He, Zefeng, et al.
Published: (2025)
VarAD: Lightweight High-Resolution Image Anomaly Detection via Visual Autoregressive Modeling
by: Cao, Yunkang, et al.
Published: (2024)
by: Cao, Yunkang, et al.
Published: (2024)
DressCode: Autoregressively Sewing and Generating Garments from Text Guidance
by: He, Kai, et al.
Published: (2024)
by: He, Kai, et al.
Published: (2024)
Dense Cross-Scale Image Alignment With Fully Spatial Correlation and Just Noticeable Difference Guidance
by: You, Jinkun, et al.
Published: (2025)
by: You, Jinkun, et al.
Published: (2025)
Extrapolating and Decoupling Image-to-Video Generation Models: Motion Modeling is Easier Than You Think
by: Tian, Jie, et al.
Published: (2025)
by: Tian, Jie, et al.
Published: (2025)
Persistent Visual Memory: Sustaining Perception for Deep Generation in LVLMs
by: Huang, Siyuan, et al.
Published: (2026)
by: Huang, Siyuan, et al.
Published: (2026)
Visual Autoregressive Modeling for Image Super-Resolution
by: Qu, Yunpeng, et al.
Published: (2025)
by: Qu, Yunpeng, et al.
Published: (2025)
AID: Attention Interpolation of Text-to-Image Diffusion
by: He, Qiyuan, et al.
Published: (2024)
by: He, Qiyuan, et al.
Published: (2024)
Autoregressive Image Generation without Vector Quantization
by: Li, Tianhong, et al.
Published: (2024)
by: Li, Tianhong, et al.
Published: (2024)
Generalizing Vision-Language Models with Dedicated Prompt Guidance
by: Li, Xinyao, et al.
Published: (2025)
by: Li, Xinyao, et al.
Published: (2025)
Randomized Autoregressive Visual Generation
by: Yu, Qihang, et al.
Published: (2024)
by: Yu, Qihang, et al.
Published: (2024)
From Sequential to Spatial: Reordering Autoregression for Efficient Visual Generation
by: Wang, Siyang, et al.
Published: (2025)
by: Wang, Siyang, et al.
Published: (2025)
DiffThinker: Towards Generative Multimodal Reasoning with Diffusion Models
by: He, Zefeng, et al.
Published: (2025)
by: He, Zefeng, et al.
Published: (2025)
InstructHumans: Editing Animated 3D Human Textures with Instructions
by: Zhu, Jiayin, et al.
Published: (2024)
by: Zhu, Jiayin, et al.
Published: (2024)
Bridging Continuous and Discrete Tokens for Autoregressive Visual Generation
by: Wang, Yuqing, et al.
Published: (2025)
by: Wang, Yuqing, et al.
Published: (2025)
Skywork UniPic: Unified Autoregressive Modeling for Visual Understanding and Generation
by: Wang, Peiyu, et al.
Published: (2025)
by: Wang, Peiyu, et al.
Published: (2025)
Causality Model for Semantic Understanding on Videos
by: Yicong, Li
Published: (2025)
by: Yicong, Li
Published: (2025)
CAR: Controllable Autoregressive Modeling for Visual Generation
by: Yao, Ziyu, et al.
Published: (2024)
by: Yao, Ziyu, et al.
Published: (2024)
Visual Intention Grounding for Egocentric Assistants
by: Sun, Pengzhan, et al.
Published: (2025)
by: Sun, Pengzhan, et al.
Published: (2025)
FrameThinker: Learning to Think with Long Videos via Multi-Turn Frame Spotlighting
by: He, Zefeng, et al.
Published: (2025)
by: He, Zefeng, et al.
Published: (2025)
ARVideo: Autoregressive Pretraining for Self-Supervised Video Representation Learning
by: Ren, Sucheng, et al.
Published: (2024)
by: Ren, Sucheng, et al.
Published: (2024)
Progressive Autoregressive Video Diffusion Models
by: Xie, Desai, et al.
Published: (2024)
by: Xie, Desai, et al.
Published: (2024)
Autoregressive Image Generation with Vision Full-view Prompt
by: Cai, Miaomiao, et al.
Published: (2025)
by: Cai, Miaomiao, et al.
Published: (2025)
SoftCFG: Uncertainty-guided Stable Guidance for Visual Autoregressive Model
by: Xu, Dongli, et al.
Published: (2025)
by: Xu, Dongli, et al.
Published: (2025)
Similar Items
-
VA-$π$: Variational Policy Alignment for Pixel-Aware Autoregressive Generation
by: Liao, Xinyao, et al.
Published: (2025) -
REAR: Rethinking Visual Autoregressive Models via Generator-Tokenizer Consistency Regularization
by: He, Qiyuan, et al.
Published: (2025) -
Interp3D: Correspondence-aware Interpolation for Generative Textured 3D Morphing
by: Liu, Xiaolu, et al.
Published: (2026) -
RelaxFlow: Text-Driven Amodal 3D Generation
by: Zhu, Jiayin, et al.
Published: (2026) -
Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning
by: Liao, Xinyao, et al.
Published: (2025)