Next Patch Prediction for Autoregressive Visual Generation
Fuente:
arXiv
Salvato in:
| Autori principali: | Pang, Yatian, Jin, Peng, Yang, Shuo, Lin, Bin, Zhu, Bin, Tang, Zhenyu, Chen, Liuhan, Tay, Francis E. H., Lim, Ser-Nam, Yang, Harry, Yuan, Li |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
DreamDance: Animating Human Images by Enriching 3D Geometry Cues from 2D Poses
di: Pang, Yatian, et al.
Pubblicazione: (2024)
di: Pang, Yatian, et al.
Pubblicazione: (2024)
VideoMerge: Towards Training-free Long Video Generation
di: Zhang, Siyang, et al.
Pubblicazione: (2025)
di: Zhang, Siyang, et al.
Pubblicazione: (2025)
Beyond Generation: Unlocking Universal Editing via Self-Supervised Fine-Tuning
di: Chen, Harold Haodong, et al.
Pubblicazione: (2024)
di: Chen, Harold Haodong, et al.
Pubblicazione: (2024)
Cycle3D: High-quality and Consistent Image-to-3D Generation via Generation-Reconstruction Cycle
di: Tang, Zhenyu, et al.
Pubblicazione: (2024)
di: Tang, Zhenyu, et al.
Pubblicazione: (2024)
MoE-LLaVA: Mixture of Experts for Large Vision-Language Models
di: Lin, Bin, et al.
Pubblicazione: (2024)
di: Lin, Bin, et al.
Pubblicazione: (2024)
VideoGen-of-Thought: Step-by-step generating multi-shot video with minimal manual intervention
di: Zheng, Mingzhe, et al.
Pubblicazione: (2025)
di: Zheng, Mingzhe, et al.
Pubblicazione: (2025)
VideoGen-of-Thought: Step-by-step generating multi-shot video with minimal manual intervention
di: Zheng, Mingzhe, et al.
Pubblicazione: (2024)
di: Zheng, Mingzhe, et al.
Pubblicazione: (2024)
Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction
di: Tian, Keyu, et al.
Pubblicazione: (2024)
di: Tian, Keyu, et al.
Pubblicazione: (2024)
Towards Chunk-Wise Generation for Long Videos
di: Zhang, Siyang, et al.
Pubblicazione: (2024)
di: Zhang, Siyang, et al.
Pubblicazione: (2024)
Object Recognition as Next Token Prediction
di: Yue, Kaiyu, et al.
Pubblicazione: (2023)
di: Yue, Kaiyu, et al.
Pubblicazione: (2023)
Hierarchical Fine-grained Preference Optimization for Physically Plausible Video Generation
di: Chen, Harold Haodong, et al.
Pubblicazione: (2025)
di: Chen, Harold Haodong, et al.
Pubblicazione: (2025)
Envision3D: One Image to 3D with Anchor Views Interpolation
di: Pang, Yatian, et al.
Pubblicazione: (2024)
di: Pang, Yatian, et al.
Pubblicazione: (2024)
Open-Sora Plan: Open-Source Large Video Generation Model
di: Lin, Bin, et al.
Pubblicazione: (2024)
di: Lin, Bin, et al.
Pubblicazione: (2024)
SwapAnyone: Consistent and Realistic Video Synthesis for Swapping Any Person into Any Video
di: Zhao, Chengshu, et al.
Pubblicazione: (2025)
di: Zhao, Chengshu, et al.
Pubblicazione: (2025)
Beyond Next-Token: Next-X Prediction for Autoregressive Visual Generation
di: Ren, Sucheng, et al.
Pubblicazione: (2025)
di: Ren, Sucheng, et al.
Pubblicazione: (2025)
Autoregressive Video Generation beyond Next Frames Prediction
di: Ren, Sucheng, et al.
Pubblicazione: (2025)
di: Ren, Sucheng, et al.
Pubblicazione: (2025)
WF-VAE: Enhancing Video VAE by Wavelet-Driven Energy Flow for Latent Video Diffusion Model
di: Li, Zongjian, et al.
Pubblicazione: (2024)
di: Li, Zongjian, et al.
Pubblicazione: (2024)
UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation
di: Lin, Bin, et al.
Pubblicazione: (2025)
di: Lin, Bin, et al.
Pubblicazione: (2025)
Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
di: Lin, Bin, et al.
Pubblicazione: (2023)
di: Lin, Bin, et al.
Pubblicazione: (2023)
AC-Foley: Reference-Audio-Guided Video-to-Audio Synthesis with Acoustic Transfer
di: Fang, Pengjun, et al.
Pubblicazione: (2026)
di: Fang, Pengjun, et al.
Pubblicazione: (2026)
Niagara: Normal-Integrated Geometric Affine Field for Scene Reconstruction from a Single View
di: Wu, Xianzu, et al.
Pubblicazione: (2025)
di: Wu, Xianzu, et al.
Pubblicazione: (2025)
Zero-shot Synthetic Video Realism Enhancement via Structure-aware Denoising
di: Wang, Yifan, et al.
Pubblicazione: (2025)
di: Wang, Yifan, et al.
Pubblicazione: (2025)
OD-VAE: An Omni-dimensional Video Compressor for Improving Latent Video Diffusion Model
di: Chen, Liuhan, et al.
Pubblicazione: (2024)
di: Chen, Liuhan, et al.
Pubblicazione: (2024)
Visual Delta Generator with Large Multi-modal Models for Semi-supervised Composed Image Retrieval
di: Jang, Young Kyun, et al.
Pubblicazione: (2024)
di: Jang, Young Kyun, et al.
Pubblicazione: (2024)
AirSketch: Generative Motion to Sketch
di: Lim, Hui Xian Grace, et al.
Pubblicazione: (2024)
di: Lim, Hui Xian Grace, et al.
Pubblicazione: (2024)
Look-Back: Implicit Visual Re-focusing in MLLM Reasoning
di: Yang, Shuo, et al.
Pubblicazione: (2025)
di: Yang, Shuo, et al.
Pubblicazione: (2025)
Learning from Next-Frame Prediction: Autoregressive Video Modeling Encodes Effective Representations
di: Li, Jinghan, et al.
Pubblicazione: (2025)
di: Li, Jinghan, et al.
Pubblicazione: (2025)
FVAR: Visual Autoregressive Modeling via Next Focus Prediction
di: Li, Xiaofan, et al.
Pubblicazione: (2025)
di: Li, Xiaofan, et al.
Pubblicazione: (2025)
Scene Co-pilot: Procedural Text to Video Generation with Human in the Loop
di: Qian, Zhaofang, et al.
Pubblicazione: (2024)
di: Qian, Zhaofang, et al.
Pubblicazione: (2024)
Video Decomposition Prior: A Methodology to Decompose Videos into Layers
di: Shrivastava, Gaurav, et al.
Pubblicazione: (2024)
di: Shrivastava, Gaurav, et al.
Pubblicazione: (2024)
GenAR: Next-Scale Autoregressive Generation for Spatial Gene Expression Prediction
di: Ouyang, Jiarui, et al.
Pubblicazione: (2025)
di: Ouyang, Jiarui, et al.
Pubblicazione: (2025)
AlignVid: Training-Free Attention Scaling for Semantic Fidelity in Text-Guided Image-to-Video Generation
di: Liu, Yexin, et al.
Pubblicazione: (2025)
di: Liu, Yexin, et al.
Pubblicazione: (2025)
Trajeglish: Traffic Modeling as Next-Token Prediction
di: Philion, Jonah, et al.
Pubblicazione: (2023)
di: Philion, Jonah, et al.
Pubblicazione: (2023)
NARRepair: Non-Autoregressive Code Generation Model for Automatic Program Repair
di: Yang, Zhenyu, et al.
Pubblicazione: (2024)
di: Yang, Zhenyu, et al.
Pubblicazione: (2024)
BOOKAGENT: Orchestrating Safety-Aware Visual Narratives via Multi-Agent Cognitive Calibration
di: Gao, Bo, et al.
Pubblicazione: (2026)
di: Gao, Bo, et al.
Pubblicazione: (2026)
Temporal Regularization Makes Your Video Generator Stronger
di: Chen, Harold Haodong, et al.
Pubblicazione: (2025)
di: Chen, Harold Haodong, et al.
Pubblicazione: (2025)
Is This Predictor More Informative than Another? A Decision-Theoretical Comparison
di: Feng, Yiding, et al.
Pubblicazione: (2025)
di: Feng, Yiding, et al.
Pubblicazione: (2025)
Parallelized Autoregressive Visual Generation
di: Wang, Yuqing, et al.
Pubblicazione: (2024)
di: Wang, Yuqing, et al.
Pubblicazione: (2024)
DiReCT: Disentangled Regularization of Contrastive Trajectories for Physics-Refined Video Generation
di: Meyarian, Abolfazl, et al.
Pubblicazione: (2026)
di: Meyarian, Abolfazl, et al.
Pubblicazione: (2026)
OSP-Next: Efficient High-Quality Video Generation with Sparse Sequence Parallelism, HiF8 Quantization, and Reinforcement Learning
di: Ge, Yunyang, et al.
Pubblicazione: (2026)
di: Ge, Yunyang, et al.
Pubblicazione: (2026)
Documenti analoghi
-
DreamDance: Animating Human Images by Enriching 3D Geometry Cues from 2D Poses
di: Pang, Yatian, et al.
Pubblicazione: (2024) -
VideoMerge: Towards Training-free Long Video Generation
di: Zhang, Siyang, et al.
Pubblicazione: (2025) -
Beyond Generation: Unlocking Universal Editing via Self-Supervised Fine-Tuning
di: Chen, Harold Haodong, et al.
Pubblicazione: (2024) -
Cycle3D: High-quality and Consistent Image-to-3D Generation via Generation-Reconstruction Cycle
di: Tang, Zhenyu, et al.
Pubblicazione: (2024) -
MoE-LLaVA: Mixture of Experts for Large Vision-Language Models
di: Lin, Bin, et al.
Pubblicazione: (2024)