VideoVerse: Does Your T2V Generator Have World Model Capability to Synthesize Videos?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Zeqing, Wei, Xinyu, Li, Bairui, Guo, Zhen, Zhang, Jinrui, Wei, Hongyang, Wang, Keze, Zhang, Lei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
TIIF-Bench: How Does Your T2I Model Follow Your Instructions?
von: Wei, Xinyu, et al.
Veröffentlicht: (2025)
von: Wei, Xinyu, et al.
Veröffentlicht: (2025)
MICo-150K: A Comprehensive Dataset Advancing Multi-Image Composition
von: Wei, Xinyu, et al.
Veröffentlicht: (2025)
von: Wei, Xinyu, et al.
Veröffentlicht: (2025)
PhyDetEx: Detecting and Explaining the Physical Plausibility of T2V Models
von: Wang, Zeqing, et al.
Veröffentlicht: (2025)
von: Wang, Zeqing, et al.
Veröffentlicht: (2025)
ShareVerse: Multi-Agent Consistent Video Generation for Shared World Modeling
von: Zhu, Jiayi, et al.
Veröffentlicht: (2026)
von: Zhu, Jiayi, et al.
Veröffentlicht: (2026)
NeoVerse: Enhancing 4D World Model with in-the-wild Monocular Videos
von: Yang, Yuxue, et al.
Veröffentlicht: (2026)
von: Yang, Yuxue, et al.
Veröffentlicht: (2026)
TimeCausality: Evaluating the Causal Ability in Time Dimension for Vision Language Models
von: Wang, Zeqing, et al.
Veröffentlicht: (2025)
von: Wang, Zeqing, et al.
Veröffentlicht: (2025)
DeepVerse: 4D Autoregressive Video Generation as a World Model
von: Chen, Junyi, et al.
Veröffentlicht: (2025)
von: Chen, Junyi, et al.
Veröffentlicht: (2025)
ALIVE: Animate Your World with Lifelike Audio-Video Generation
von: Guo, Ying, et al.
Veröffentlicht: (2026)
von: Guo, Ying, et al.
Veröffentlicht: (2026)
EditVerse: Unifying Image and Video Editing and Generation with In-Context Learning
von: Ju, Xuan, et al.
Veröffentlicht: (2025)
von: Ju, Xuan, et al.
Veröffentlicht: (2025)
UniVerse-1: Unified Audio-Video Generation via Stitching of Experts
von: Wang, Duomin, et al.
Veröffentlicht: (2025)
von: Wang, Duomin, et al.
Veröffentlicht: (2025)
MultiCounter: Multiple Action Agnostic Repetition Counting in Untrimmed Videos
von: Tang, Yin, et al.
Veröffentlicht: (2024)
von: Tang, Yin, et al.
Veröffentlicht: (2024)
Let Your Video Listen to Your Music!
von: Zhang, Xinyu, et al.
Veröffentlicht: (2025)
von: Zhang, Xinyu, et al.
Veröffentlicht: (2025)
DriveWAM: Video Generative Priors Enable Scalable World-Action Modeling for Autonomous Driving
von: Shi, Chen, et al.
Veröffentlicht: (2026)
von: Shi, Chen, et al.
Veröffentlicht: (2026)
How to Design and Train Your Implicit Neural Representation for Video Compression
von: Gwilliam, Matthew, et al.
Veröffentlicht: (2025)
von: Gwilliam, Matthew, et al.
Veröffentlicht: (2025)
Perceive, Understand and Restore: Real-World Image Super-Resolution with Autoregressive Multimodal Generative Models
von: Wei, Hongyang, et al.
Veröffentlicht: (2025)
von: Wei, Hongyang, et al.
Veröffentlicht: (2025)
InstaDrive: Instance-Aware Driving World Models for Realistic and Consistent Video Generation
von: Yang, Zhuoran, et al.
Veröffentlicht: (2026)
von: Yang, Zhuoran, et al.
Veröffentlicht: (2026)
MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?
von: Zhang, Renrui, et al.
Veröffentlicht: (2024)
von: Zhang, Renrui, et al.
Veröffentlicht: (2024)
Seeing More, Saying More: Lightweight Language Experts are Dynamic Video Token Compressors
von: Wang, Xiangchen, et al.
Veröffentlicht: (2025)
von: Wang, Xiangchen, et al.
Veröffentlicht: (2025)
EvalVerse: Pipeline-Aware and Expert-Calibrated Benchmarking for Professional Cinematic Video Generation
von: Yang, Songlin, et al.
Veröffentlicht: (2026)
von: Yang, Songlin, et al.
Veröffentlicht: (2026)
CascadeV: An Implementation of Wurstchen Architecture for Video Generation
von: Lin, Wenfeng, et al.
Veröffentlicht: (2025)
von: Lin, Wenfeng, et al.
Veröffentlicht: (2025)
MagicInfinite: Generating Infinite Talking Videos with Your Words and Voice
von: Yi, Hongwei, et al.
Veröffentlicht: (2025)
von: Yi, Hongwei, et al.
Veröffentlicht: (2025)
STORM: Search-Guided Generative World Models for Robotic Manipulation
von: Lin, Wenjun, et al.
Veröffentlicht: (2025)
von: Lin, Wenjun, et al.
Veröffentlicht: (2025)
TalkVerse: Democratizing Minute-Long Audio-Driven Video Generation
von: Wang, Zhenzhi, et al.
Veröffentlicht: (2025)
von: Wang, Zhenzhi, et al.
Veröffentlicht: (2025)
Process-of-Thought Reasoning for Videos
von: Zhang, Jusheng, et al.
Veröffentlicht: (2026)
von: Zhang, Jusheng, et al.
Veröffentlicht: (2026)
Is this Generated Person Existed in Real-world? Fine-grained Detecting and Calibrating Abnormal Human-body
von: Wang, Zeqing, et al.
Veröffentlicht: (2024)
von: Wang, Zeqing, et al.
Veröffentlicht: (2024)
MBench: A Comprehensive Benchmark on Memory Capability for Video World Models
von: Zhang, Shengjun, et al.
Veröffentlicht: (2026)
von: Zhang, Shengjun, et al.
Veröffentlicht: (2026)
LLMVA-GEBC: Large Language Model with Video Adapter for Generic Event Boundary Captioning
von: Tang, Yolo Yunlong, et al.
Veröffentlicht: (2023)
von: Tang, Yolo Yunlong, et al.
Veröffentlicht: (2023)
Spatia: Video Generation with Updatable Spatial Memory
von: Zhao, Jinjing, et al.
Veröffentlicht: (2025)
von: Zhao, Jinjing, et al.
Veröffentlicht: (2025)
Multi-sentence Video Grounding for Long Video Generation
von: Feng, Wei, et al.
Veröffentlicht: (2024)
von: Feng, Wei, et al.
Veröffentlicht: (2024)
VerseCrafter: Dynamic Realistic Video World Model with 4D Geometric Control
von: Zheng, Sixiao, et al.
Veröffentlicht: (2026)
von: Zheng, Sixiao, et al.
Veröffentlicht: (2026)
MechVerse: Evaluating Physical Motion Consistency in Video Generation Models
von: Jain, Rahul, et al.
Veröffentlicht: (2026)
von: Jain, Rahul, et al.
Veröffentlicht: (2026)
WonderVerse: Extendable 3D Scene Generation with Video Generative Models
von: Feng, Hao, et al.
Veröffentlicht: (2025)
von: Feng, Hao, et al.
Veröffentlicht: (2025)
CoAgent: Collaborative Planning and Consistency Agent for Coherent Video Generation
von: Zeng, Qinglin, et al.
Veröffentlicht: (2025)
von: Zeng, Qinglin, et al.
Veröffentlicht: (2025)
Video Super-Resolution Transformer with Masked Inter&Intra-Frame Attention
von: Zhou, Xingyu, et al.
Veröffentlicht: (2024)
von: Zhou, Xingyu, et al.
Veröffentlicht: (2024)
Minute-Long Videos with Dual Parallelisms
von: Wang, Zeqing, et al.
Veröffentlicht: (2025)
von: Wang, Zeqing, et al.
Veröffentlicht: (2025)
LiveWorld: Simulating Out-of-Sight Dynamics in Generative Video World Models
von: Duan, Zicheng, et al.
Veröffentlicht: (2026)
von: Duan, Zicheng, et al.
Veröffentlicht: (2026)
StyleMaster: Stylize Your Video with Artistic Generation and Translation
von: Ye, Zixuan, et al.
Veröffentlicht: (2024)
von: Ye, Zixuan, et al.
Veröffentlicht: (2024)
ShotVerse: Advancing Cinematic Camera Control for Text-Driven Multi-Shot Video Creation
von: Yang, Songlin, et al.
Veröffentlicht: (2026)
von: Yang, Songlin, et al.
Veröffentlicht: (2026)
ModaVerse: Efficiently Transforming Modalities with LLMs
von: Wang, Xinyu, et al.
Veröffentlicht: (2024)
von: Wang, Xinyu, et al.
Veröffentlicht: (2024)
Incantation: Natural Language as the Action Interface for Multi-Entity Video World Models
von: Zhu, Shangwen, et al.
Veröffentlicht: (2026)
von: Zhu, Shangwen, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
TIIF-Bench: How Does Your T2I Model Follow Your Instructions?
von: Wei, Xinyu, et al.
Veröffentlicht: (2025) -
MICo-150K: A Comprehensive Dataset Advancing Multi-Image Composition
von: Wei, Xinyu, et al.
Veröffentlicht: (2025) -
PhyDetEx: Detecting and Explaining the Physical Plausibility of T2V Models
von: Wang, Zeqing, et al.
Veröffentlicht: (2025) -
ShareVerse: Multi-Agent Consistent Video Generation for Shared World Modeling
von: Zhu, Jiayi, et al.
Veröffentlicht: (2026) -
NeoVerse: Enhancing 4D World Model with in-the-wild Monocular Videos
von: Yang, Yuxue, et al.
Veröffentlicht: (2026)