What Happens Next? Next Scene Prediction with a Unified Video Model
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Xinjie, Chen, Zhimin, Zhao, Rui, Schiffers, Florian, Liao, Zhenyu, Bhat, Vimal |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Motion Modes: What Could Happen Next?
by: Pandey, Karran, et al.
Published: (2024)
by: Pandey, Karran, et al.
Published: (2024)
Generating Multimodal Driving Scenes via Next-Scene Prediction
by: Wu, Yanhao, et al.
Published: (2025)
by: Wu, Yanhao, et al.
Published: (2025)
Video-as-Answer: Predict and Generate Next Video Event with Joint-GRPO
by: Cheng, Junhao, et al.
Published: (2025)
by: Cheng, Junhao, et al.
Published: (2025)
What Happens Next? Anticipating Future Motion by Generating Point Trajectories
by: Boduljak, Gabrijel, et al.
Published: (2025)
by: Boduljak, Gabrijel, et al.
Published: (2025)
Masked Next-Scale Prediction for Self-supervised Scene Text Recognition
by: Chen, Zhuohao, et al.
Published: (2026)
by: Chen, Zhuohao, et al.
Published: (2026)
Autoregressive Video Generation beyond Next Frames Prediction
by: Ren, Sucheng, et al.
Published: (2025)
by: Ren, Sucheng, et al.
Published: (2025)
VeRVE: Versatile Retrieval for Videos via Unified Embeddings
by: Halbe, Shaunak, et al.
Published: (2026)
by: Halbe, Shaunak, et al.
Published: (2026)
Text-Guided Video Masked Autoencoder
by: Fan, David, et al.
Published: (2024)
by: Fan, David, et al.
Published: (2024)
DiffSign: AI-Assisted Generation of Customizable Sign Language Videos With Enhanced Realism
by: Krishnamurthy, Sudha, et al.
Published: (2024)
by: Krishnamurthy, Sudha, et al.
Published: (2024)
Everything is a Video: Unifying Modalities through Next-Frame Prediction
by: Hudson, G. Thomas, et al.
Published: (2024)
by: Hudson, G. Thomas, et al.
Published: (2024)
Long-Context Autoregressive Video Modeling with Next-Frame Prediction
by: Gu, Yuchao, et al.
Published: (2025)
by: Gu, Yuchao, et al.
Published: (2025)
Next Patch Prediction for Autoregressive Visual Generation
by: Pang, Yatian, et al.
Published: (2024)
by: Pang, Yatian, et al.
Published: (2024)
Beyond Next-Token: Next-X Prediction for Autoregressive Visual Generation
by: Ren, Sucheng, et al.
Published: (2025)
by: Ren, Sucheng, et al.
Published: (2025)
Frame Context Packing and Drift Prevention in Next-Frame-Prediction Video Diffusion Models
by: Zhang, Lvmin, et al.
Published: (2025)
by: Zhang, Lvmin, et al.
Published: (2025)
UniVA: Universal Video Agent towards Open-Source Next-Generation Video Generalist
by: Liang, Zhengyang, et al.
Published: (2025)
by: Liang, Zhengyang, et al.
Published: (2025)
Learning from Next-Frame Prediction: Autoregressive Video Modeling Encodes Effective Representations
by: Li, Jinghan, et al.
Published: (2025)
by: Li, Jinghan, et al.
Published: (2025)
Object Recognition as Next Token Prediction
by: Yue, Kaiyu, et al.
Published: (2023)
by: Yue, Kaiyu, et al.
Published: (2023)
SceneStreamer: Continuous Scenario Generation as Next Token Group Prediction
by: Peng, Zhenghao, et al.
Published: (2025)
by: Peng, Zhenghao, et al.
Published: (2025)
AMP: Autoregressive Motion Prediction Revisited with Next Token Prediction for Autonomous Driving
by: Jia, Xiaosong, et al.
Published: (2024)
by: Jia, Xiaosong, et al.
Published: (2024)
NEP: Autoregressive Image Editing via Next Editing Token Prediction
by: Wu, Huimin, et al.
Published: (2025)
by: Wu, Huimin, et al.
Published: (2025)
STCOcc: Sparse Spatial-Temporal Cascade Renovation for 3D Occupancy and Scene Flow Prediction
by: Liao, Zhimin, et al.
Published: (2025)
by: Liao, Zhimin, et al.
Published: (2025)
Detect Anything via Next Point Prediction
by: Jiang, Qing, et al.
Published: (2025)
by: Jiang, Qing, et al.
Published: (2025)
NextFlow: Unified Sequential Modeling Activates Multimodal Understanding and Generation
by: Zhang, Huichao, et al.
Published: (2026)
by: Zhang, Huichao, et al.
Published: (2026)
Next Block Prediction: Video Generation via Semi-Autoregressive Modeling
by: Ren, Shuhuai, et al.
Published: (2025)
by: Ren, Shuhuai, et al.
Published: (2025)
TempoMaster: Efficient Long Video Generation via Next-Frame-Rate Prediction
by: Ma, Yukuo, et al.
Published: (2025)
by: Ma, Yukuo, et al.
Published: (2025)
SGG-R$^{\rm 3}$: From Next-Token Prediction to End-to-End Unbiased Scene Graph Generation
by: Feng, Jiaye, et al.
Published: (2026)
by: Feng, Jiaye, et al.
Published: (2026)
FVAR: Visual Autoregressive Modeling via Next Focus Prediction
by: Li, Xiaofan, et al.
Published: (2025)
by: Li, Xiaofan, et al.
Published: (2025)
NextAds: Towards Next-generation Personalized Video Advertising
by: Xu, Yiyan, et al.
Published: (2026)
by: Xu, Yiyan, et al.
Published: (2026)
AnimeGamer: Infinite Anime Life Simulation with Next Game State Prediction
by: Cheng, Junhao, et al.
Published: (2025)
by: Cheng, Junhao, et al.
Published: (2025)
Anticipating Next Active Objects for Egocentric Videos
by: Thakur, Sanket, et al.
Published: (2023)
by: Thakur, Sanket, et al.
Published: (2023)
TAPNext++: What's Next for Tracking Any Point (TAP)?
by: Jung, Sebastian, et al.
Published: (2026)
by: Jung, Sebastian, et al.
Published: (2026)
Lumina-Video: Efficient and Flexible Video Generation with Multi-scale Next-DiT
by: Liu, Dongyang, et al.
Published: (2025)
by: Liu, Dongyang, et al.
Published: (2025)
SAM-Guided Masked Token Prediction for 3D Scene Understanding
by: Chen, Zhimin, et al.
Published: (2024)
by: Chen, Zhimin, et al.
Published: (2024)
Fostering Video Reasoning via Next-Event Prediction
by: Wang, Haonan, et al.
Published: (2025)
by: Wang, Haonan, et al.
Published: (2025)
A Simple Baseline for Unifying Understanding, Generation, and Editing via Vanilla Next-token Prediction
by: Zhu, Jie, et al.
Published: (2026)
by: Zhu, Jie, et al.
Published: (2026)
What Happens When: Learning Temporal Orders of Events in Videos
by: Ahn, Daechul, et al.
Published: (2025)
by: Ahn, Daechul, et al.
Published: (2025)
InternVideo-Next: Towards General Video Foundation Models without Video-Text Supervision
by: Wang, Chenting, et al.
Published: (2025)
by: Wang, Chenting, et al.
Published: (2025)
Emu3: Next-Token Prediction is All You Need
by: Wang, Xinlong, et al.
Published: (2024)
by: Wang, Xinlong, et al.
Published: (2024)
EgoIntent: An Egocentric Step-level Benchmark for Understanding What, Why, and Next
by: Pan, Ye, et al.
Published: (2026)
by: Pan, Ye, et al.
Published: (2026)
Next-Embedding Prediction Makes Strong Vision Learners
by: Xu, Sihan, et al.
Published: (2025)
by: Xu, Sihan, et al.
Published: (2025)
Similar Items
-
Motion Modes: What Could Happen Next?
by: Pandey, Karran, et al.
Published: (2024) -
Generating Multimodal Driving Scenes via Next-Scene Prediction
by: Wu, Yanhao, et al.
Published: (2025) -
Video-as-Answer: Predict and Generate Next Video Event with Joint-GRPO
by: Cheng, Junhao, et al.
Published: (2025) -
What Happens Next? Anticipating Future Motion by Generating Point Trajectories
by: Boduljak, Gabrijel, et al.
Published: (2025) -
Masked Next-Scale Prediction for Self-supervised Scene Text Recognition
by: Chen, Zhuohao, et al.
Published: (2026)