DeepVerse: 4D Autoregressive Video Generation as a World Model
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Chen, Junyi, Zhu, Haoyi, He, Xianglong, Wang, Yifan, Zhou, Jianjun, Chang, Wenzheng, Zhou, Yang, Li, Zizun, Fu, Zhoujie, Pang, Jiangmiao, He, Tong |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Aether: Geometric-Aware Unified World Modeling
par: Aether Team, et autres
Publié: (2025)
par: Aether Team, et autres
Publié: (2025)
$π^3$: Permutation-Equivariant Visual Geometry Learning
par: Wang, Yifan, et autres
Publié: (2025)
par: Wang, Yifan, et autres
Publié: (2025)
OmniWorld: A Multi-Domain and Multi-Modal Dataset for 4D World Modeling
par: Zhou, Yang, et autres
Publié: (2025)
par: Zhou, Yang, et autres
Publié: (2025)
WinT3R: Window-Based Streaming Reconstruction with Camera Token Pool
par: Li, Zizun, et autres
Publié: (2025)
par: Li, Zizun, et autres
Publié: (2025)
Geo-Align: Video Generation Alignment via Metric Geometry Reward
par: Li, Zizun, et autres
Publié: (2026)
par: Li, Zizun, et autres
Publié: (2026)
VINO: A Unified Visual Generator with Interleaved OmniModal Context
par: Chen, Junyi, et autres
Publié: (2026)
par: Chen, Junyi, et autres
Publié: (2026)
AR4D: Autoregressive 4D Generation from Monocular Videos
par: Zhu, Hanxin, et autres
Publié: (2025)
par: Zhu, Hanxin, et autres
Publié: (2025)
VerseCrafter: Dynamic Realistic Video World Model with 4D Geometric Control
par: Zheng, Sixiao, et autres
Publié: (2026)
par: Zheng, Sixiao, et autres
Publié: (2026)
Warp-as-History: Generalizable Camera-Controlled Video Generation from One Training Video
par: Wang, Yifan, et autres
Publié: (2026)
par: Wang, Yifan, et autres
Publié: (2026)
UrbanVerse: Scaling Urban Simulation by Watching City-Tour Videos
par: Liu, Mingxuan, et autres
Publié: (2025)
par: Liu, Mingxuan, et autres
Publié: (2025)
NeoVerse: Enhancing 4D World Model with in-the-wild Monocular Videos
par: Yang, Yuxue, et autres
Publié: (2026)
par: Yang, Yuxue, et autres
Publié: (2026)
ShareVerse: Multi-Agent Consistent Video Generation for Shared World Modeling
par: Zhu, Jiayi, et autres
Publié: (2026)
par: Zhu, Jiayi, et autres
Publié: (2026)
GVGEN: Text-to-3D Generation with Volumetric Representation
par: He, Xianglong, et autres
Publié: (2024)
par: He, Xianglong, et autres
Publié: (2024)
Sync4D: Video Guided Controllable Dynamics for Physics-Based 4D Generation
par: Fu, Zhoujie, et autres
Publié: (2024)
par: Fu, Zhoujie, et autres
Publié: (2024)
Yume: An Interactive World Generation Model
par: Mao, Xiaofeng, et autres
Publié: (2025)
par: Mao, Xiaofeng, et autres
Publié: (2025)
EditVerse: Unifying Image and Video Editing and Generation with In-Context Learning
par: Ju, Xuan, et autres
Publié: (2025)
par: Ju, Xuan, et autres
Publié: (2025)
Yume-1.5: A Text-Controlled Interactive World Generation Model
par: Mao, Xiaofeng, et autres
Publié: (2025)
par: Mao, Xiaofeng, et autres
Publié: (2025)
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs
par: He, Yuping, et autres
Publié: (2025)
par: He, Yuping, et autres
Publié: (2025)
TalkVerse: Democratizing Minute-Long Audio-Driven Video Generation
par: Wang, Zhenzhi, et autres
Publié: (2025)
par: Wang, Zhenzhi, et autres
Publié: (2025)
Neighboring Autoregressive Modeling for Efficient Visual Generation
par: He, Yefei, et autres
Publié: (2025)
par: He, Yefei, et autres
Publié: (2025)
HiTVideo: Hierarchical Tokenizers for Enhancing Text-to-Video Generation with Autoregressive Large Language Models
par: Zhou, Ziqin, et autres
Publié: (2025)
par: Zhou, Ziqin, et autres
Publié: (2025)
PonderV2: Pave the Way for 3D Foundation Model with A Universal Pre-training Paradigm
par: Zhu, Haoyi, et autres
Publié: (2023)
par: Zhu, Haoyi, et autres
Publié: (2023)
WonderVerse: Extendable 3D Scene Generation with Video Generative Models
par: Feng, Hao, et autres
Publié: (2025)
par: Feng, Hao, et autres
Publié: (2025)
World-R1: Reinforcing 3D Constraints for Text-to-Video Generation
par: Wang, Weijie, et autres
Publié: (2026)
par: Wang, Weijie, et autres
Publié: (2026)
VideoVerse: Does Your T2V Generator Have World Model Capability to Synthesize Videos?
par: Wang, Zeqing, et autres
Publié: (2025)
par: Wang, Zeqing, et autres
Publié: (2025)
Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM
par: Huang, Haifeng, et autres
Publié: (2026)
par: Huang, Haifeng, et autres
Publié: (2026)
VPG: Visual Prefix Guidance for Autoregressive Image and Video Generation
par: Liao, Xinyao, et autres
Publié: (2026)
par: Liao, Xinyao, et autres
Publié: (2026)
Astra: General Interactive World Model with Autoregressive Denoising
par: Zhu, Yixuan, et autres
Publié: (2025)
par: Zhu, Yixuan, et autres
Publié: (2025)
Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion
par: Huang, Xun, et autres
Publié: (2025)
par: Huang, Xun, et autres
Publié: (2025)
UniVerse-1: Unified Audio-Video Generation via Stitching of Experts
par: Wang, Duomin, et autres
Publié: (2025)
par: Wang, Duomin, et autres
Publié: (2025)
3D4D: An Interactive, Editable, 4D World Model via 3D Video Generation
par: He, Yunhong, et autres
Publié: (2025)
par: He, Yunhong, et autres
Publié: (2025)
DCARL: A Divide-and-Conquer Framework for Autoregressive Long-Trajectory Video Generation
par: Ouyang, Junyi, et autres
Publié: (2026)
par: Ouyang, Junyi, et autres
Publié: (2026)
Context-Aware Autoregressive Models for Multi-Conditional Image Generation
par: Chen, Yixiao, et autres
Publié: (2025)
par: Chen, Yixiao, et autres
Publié: (2025)
DynamicVerse: A Physically-Aware Multimodal Framework for 4D World Modeling
par: Wen, Kairun, et autres
Publié: (2025)
par: Wen, Kairun, et autres
Publié: (2025)
Training-Free Watermarking for Autoregressive Image Generation
par: Tong, Yu, et autres
Publié: (2025)
par: Tong, Yu, et autres
Publié: (2025)
SPA: 3D Spatial-Awareness Enables Effective Embodied Representation
par: Zhu, Haoyi, et autres
Publié: (2024)
par: Zhu, Haoyi, et autres
Publié: (2024)
Fast Autoregressive Video Generation with Diagonal Decoding
par: Ye, Yang, et autres
Publié: (2025)
par: Ye, Yang, et autres
Publié: (2025)
LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness
par: Zhu, Chenming, et autres
Publié: (2024)
par: Zhu, Chenming, et autres
Publié: (2024)
HOI4D: A 4D Egocentric Dataset for Category-Level Human-Object Interaction
par: Liu, Yunze, et autres
Publié: (2022)
par: Liu, Yunze, et autres
Publié: (2022)
ECViT: Efficient Convolutional Vision Transformer with Local-Attention and Multi-scale Stages
par: Qian, Zhoujie
Publié: (2025)
par: Qian, Zhoujie
Publié: (2025)
Documents similaires
-
Aether: Geometric-Aware Unified World Modeling
par: Aether Team, et autres
Publié: (2025) -
$π^3$: Permutation-Equivariant Visual Geometry Learning
par: Wang, Yifan, et autres
Publié: (2025) -
OmniWorld: A Multi-Domain and Multi-Modal Dataset for 4D World Modeling
par: Zhou, Yang, et autres
Publié: (2025) -
WinT3R: Window-Based Streaming Reconstruction with Camera Token Pool
par: Li, Zizun, et autres
Publié: (2025) -
Geo-Align: Video Generation Alignment via Metric Geometry Reward
par: Li, Zizun, et autres
Publié: (2026)