Learning from Next-Frame Prediction: Autoregressive Video Modeling Encodes Effective Representations
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Jinghan, Jin, Yang, Jiang, Hao, Mu, Yadong, Song, Yang, Xu, Kun |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Autoregressive Video Generation beyond Next Frames Prediction
by: Ren, Sucheng, et al.
Published: (2025)
by: Ren, Sucheng, et al.
Published: (2025)
Long-Context Autoregressive Video Modeling with Next-Frame Prediction
by: Gu, Yuchao, et al.
Published: (2025)
by: Gu, Yuchao, et al.
Published: (2025)
Pyramidal Flow Matching for Efficient Video Generative Modeling
by: Jin, Yang, et al.
Published: (2024)
by: Jin, Yang, et al.
Published: (2024)
Video-LaVIT: Unified Video-Language Pre-training with Decoupled Visual-Motional Tokenization
by: Jin, Yang, et al.
Published: (2024)
by: Jin, Yang, et al.
Published: (2024)
VideoAR: Autoregressive Video Generation via Next-Frame & Scale Prediction
by: Ji, Longbin, et al.
Published: (2026)
by: Ji, Longbin, et al.
Published: (2026)
Next Patch Prediction for Autoregressive Visual Generation
by: Pang, Yatian, et al.
Published: (2024)
by: Pang, Yatian, et al.
Published: (2024)
Frame Context Packing and Drift Prevention in Next-Frame-Prediction Video Diffusion Models
by: Zhang, Lvmin, et al.
Published: (2025)
by: Zhang, Lvmin, et al.
Published: (2025)
GenAR: Next-Scale Autoregressive Generation for Spatial Gene Expression Prediction
by: Ouyang, Jiarui, et al.
Published: (2025)
by: Ouyang, Jiarui, et al.
Published: (2025)
Next Block Prediction: Video Generation via Semi-Autoregressive Modeling
by: Ren, Shuhuai, et al.
Published: (2025)
by: Ren, Shuhuai, et al.
Published: (2025)
RectifID: Personalizing Rectified Flow with Anchored Classifier Guidance
by: Sun, Zhicheng, et al.
Published: (2024)
by: Sun, Zhicheng, et al.
Published: (2024)
Pretraining Frame Preservation for Lightweight Autoregressive Video History Embedding
by: Zhang, Lvmin, et al.
Published: (2025)
by: Zhang, Lvmin, et al.
Published: (2025)
Frame-Voyager: Learning to Query Frames for Video Large Language Models
by: Yu, Sicheng, et al.
Published: (2024)
by: Yu, Sicheng, et al.
Published: (2024)
FVAR: Visual Autoregressive Modeling via Next Focus Prediction
by: Li, Xiaofan, et al.
Published: (2025)
by: Li, Xiaofan, et al.
Published: (2025)
Learning Procedural-aware Video Representations through State-Grounded Hierarchy Unfolding
by: Zhao, Jinghan, et al.
Published: (2025)
by: Zhao, Jinghan, et al.
Published: (2025)
TempoMaster: Efficient Long Video Generation via Next-Frame-Rate Prediction
by: Ma, Yukuo, et al.
Published: (2025)
by: Ma, Yukuo, et al.
Published: (2025)
Next-Scale Autoregressive Models for Text-to-Motion Generation
by: Zheng, Zhiwei, et al.
Published: (2026)
by: Zheng, Zhiwei, et al.
Published: (2026)
Asynchronous Multimodal Video Sequence Fusion via Learning Modality-Exclusive and -Agnostic Representations
by: Yang, Dingkang, et al.
Published: (2024)
by: Yang, Dingkang, et al.
Published: (2024)
Fast Encoding and Decoding for Implicit Video Representation
by: Chen, Hao, et al.
Published: (2024)
by: Chen, Hao, et al.
Published: (2024)
Video In-context Learning: Autoregressive Transformers are Zero-Shot Video Imitators
by: Zhang, Wentao, et al.
Published: (2024)
by: Zhang, Wentao, et al.
Published: (2024)
Weakly-Supervised Affordance Grounding Guided by Part-Level Semantic Priors
by: Xu, Peiran, et al.
Published: (2025)
by: Xu, Peiran, et al.
Published: (2025)
ARVideo: Autoregressive Pretraining for Self-Supervised Video Representation Learning
by: Ren, Sucheng, et al.
Published: (2024)
by: Ren, Sucheng, et al.
Published: (2024)
AMP: Autoregressive Motion Prediction Revisited with Next Token Prediction for Autonomous Driving
by: Jia, Xiaosong, et al.
Published: (2024)
by: Jia, Xiaosong, et al.
Published: (2024)
Beyond Next-Token: Next-X Prediction for Autoregressive Visual Generation
by: Ren, Sucheng, et al.
Published: (2025)
by: Ren, Sucheng, et al.
Published: (2025)
Next-Frame Decoding for Ultra-Low-Bitrate Image Compression with Video Diffusion Priors
by: Chen, Yunuo, et al.
Published: (2026)
by: Chen, Yunuo, et al.
Published: (2026)
Progressive Autoregressive Video Diffusion Models
by: Xie, Desai, et al.
Published: (2024)
by: Xie, Desai, et al.
Published: (2024)
Making Every Frame Matter: Continuous Activity Recognition in Streaming Video via Adaptive Video Context Modeling
by: Wu, Hao, et al.
Published: (2024)
by: Wu, Hao, et al.
Published: (2024)
Learning Compact Video Representations for Efficient Long-form Video Understanding in Large Multimodal Models
by: Chen, Yuxiao, et al.
Published: (2026)
by: Chen, Yuxiao, et al.
Published: (2026)
RED: Robust Environmental Design
by: Yang, Jinghan
Published: (2024)
by: Yang, Jinghan
Published: (2024)
NEP: Autoregressive Image Editing via Next Editing Token Prediction
by: Wu, Huimin, et al.
Published: (2025)
by: Wu, Huimin, et al.
Published: (2025)
Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction
by: Tian, Keyu, et al.
Published: (2024)
by: Tian, Keyu, et al.
Published: (2024)
MedVAR: Towards Scalable and Efficient Medical Image Generation via Next-scale Autoregressive Prediction
by: He, Zhicheng, et al.
Published: (2026)
by: He, Zhicheng, et al.
Published: (2026)
What Happens Next? Next Scene Prediction with a Unified Video Model
by: Li, Xinjie, et al.
Published: (2025)
by: Li, Xinjie, et al.
Published: (2025)
Not All Frames Deserve Full Computation: Accelerating Autoregressive Video Generation via Selective Computation and Predictive Extrapolation
by: Cui, Hanshuai, et al.
Published: (2026)
by: Cui, Hanshuai, et al.
Published: (2026)
DiffSplat: Repurposing Image Diffusion Models for Scalable Gaussian Splat Generation
by: Lin, Chenguo, et al.
Published: (2025)
by: Lin, Chenguo, et al.
Published: (2025)
POD: Predictive Object Detection with Single-Frame FMCW LiDAR Point Cloud
by: Shi, Yining, et al.
Published: (2025)
by: Shi, Yining, et al.
Published: (2025)
Unified Arbitrary-Time Video Frame Interpolation and Prediction
by: Jin, Xin, et al.
Published: (2025)
by: Jin, Xin, et al.
Published: (2025)
TimeTracker: Event-based Continuous Point Tracking for Video Frame Interpolation with Non-linear Motion
by: Liu, Haoyue, et al.
Published: (2025)
by: Liu, Haoyue, et al.
Published: (2025)
Neural Assembler: Learning to Generate Fine-Grained Robotic Assembly Instructions from Multi-View Images
by: Yan, Hongyu, et al.
Published: (2024)
by: Yan, Hongyu, et al.
Published: (2024)
End-to-End Training for Autoregressive Video Diffusion via Self-Resampling
by: Guo, Yuwei, et al.
Published: (2025)
by: Guo, Yuwei, et al.
Published: (2025)
DetailFlow: 1D Coarse-to-Fine Autoregressive Image Generation via Next-Detail Prediction
by: Liu, Yiheng, et al.
Published: (2025)
by: Liu, Yiheng, et al.
Published: (2025)
Similar Items
-
Autoregressive Video Generation beyond Next Frames Prediction
by: Ren, Sucheng, et al.
Published: (2025) -
Long-Context Autoregressive Video Modeling with Next-Frame Prediction
by: Gu, Yuchao, et al.
Published: (2025) -
Pyramidal Flow Matching for Efficient Video Generative Modeling
by: Jin, Yang, et al.
Published: (2024) -
Video-LaVIT: Unified Video-Language Pre-training with Decoupled Visual-Motional Tokenization
by: Jin, Yang, et al.
Published: (2024) -
VideoAR: Autoregressive Video Generation via Next-Frame & Scale Prediction
by: Ji, Longbin, et al.
Published: (2026)