Quantitative Video World Model Evaluation for Geometric-Consistency
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Jiaxin, Pi, Yihao, Zhang, Yinling, Li, Yuheng, Zou, Xueyan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PanoWorld: Geometry-Consistent Panoramic Video World Modeling
by: Jiang, Le, et al.
Published: (2026)
by: Jiang, Le, et al.
Published: (2026)
Geometry Forcing: Marrying Video Diffusion and 3D Representation for Consistent World Modeling
by: Wu, Haoyu, et al.
Published: (2025)
by: Wu, Haoyu, et al.
Published: (2025)
ShareVerse: Multi-Agent Consistent Video Generation for Shared World Modeling
by: Zhu, Jiayi, et al.
Published: (2026)
by: Zhu, Jiayi, et al.
Published: (2026)
WorldReel: 4D Video Generation with Consistent Geometry and Motion Modeling
by: Fang, Shaoheng, et al.
Published: (2025)
by: Fang, Shaoheng, et al.
Published: (2025)
RelightVid: Temporal-Consistent Diffusion Model for Video Relighting
by: Fang, Ye, et al.
Published: (2025)
by: Fang, Ye, et al.
Published: (2025)
AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
by: Wang, Zun, et al.
Published: (2026)
by: Wang, Zun, et al.
Published: (2026)
MIND: Benchmarking Memory Consistency and Action Control in World Models
by: Ye, Yixuan, et al.
Published: (2026)
by: Ye, Yixuan, et al.
Published: (2026)
Rethinking Video Generation Model for the Embodied World
by: Deng, Yufan, et al.
Published: (2026)
by: Deng, Yufan, et al.
Published: (2026)
Sekai: A Video Dataset towards World Exploration
by: Li, Zhen, et al.
Published: (2025)
by: Li, Zhen, et al.
Published: (2025)
Large-Scale 3D Medical Image Pre-training with Geometric Context Priors
by: Wu, Linshan, et al.
Published: (2024)
by: Wu, Linshan, et al.
Published: (2024)
Boximator: Generating Rich and Controllable Motions for Video Synthesis
by: Wang, Jiawei, et al.
Published: (2024)
by: Wang, Jiawei, et al.
Published: (2024)
Beyond Pixels: Introducing Geometric-Semantic World Priors for Video-based Embodied Models via Spatio-temporal Alignment
by: Tang, Jinzhou, et al.
Published: (2025)
by: Tang, Jinzhou, et al.
Published: (2025)
Mirage2Matter: A Physically Grounded Gaussian World Model from Video
by: Gao, Zhengqing, et al.
Published: (2026)
by: Gao, Zhengqing, et al.
Published: (2026)
Owl-1: Omni World Model for Consistent Long Video Generation
by: Huang, Yuanhui, et al.
Published: (2024)
by: Huang, Yuanhui, et al.
Published: (2024)
WorldGPT: A Sora-Inspired Video AI Agent as Rich World Models from Text and Image Inputs
by: Yang, Deshun, et al.
Published: (2024)
by: Yang, Deshun, et al.
Published: (2024)
Other Vehicle Trajectories Are Also Needed: A Driving World Model Unifies Ego-Other Vehicle Trajectories in Video Latent Space
by: Zhu, Jian, et al.
Published: (2025)
by: Zhu, Jian, et al.
Published: (2025)
Video Generation with Consistency Tuning
by: Wang, Chaoyi, et al.
Published: (2024)
by: Wang, Chaoyi, et al.
Published: (2024)
Self-Evolving Spatial Reasoning in Vision Language Models via Geometric Logic Consistency
by: Liu, Junming, et al.
Published: (2026)
by: Liu, Junming, et al.
Published: (2026)
WorldModelBench: Judging Video Generation Models As World Models
by: Li, Dacheng, et al.
Published: (2025)
by: Li, Dacheng, et al.
Published: (2025)
A Survey: Spatiotemporal Consistency in Video Generation
by: Yin, Zhiyu, et al.
Published: (2025)
by: Yin, Zhiyu, et al.
Published: (2025)
Long-CODE: Isolating Pure Long-Context as an Orthogonal Dimension in Video Evaluation
by: Tang, Zhijiang, et al.
Published: (2026)
by: Tang, Zhijiang, et al.
Published: (2026)
LoopNav: Benchmarking Spatial Consistency in World Models
by: Lian, Kewei, et al.
Published: (2025)
by: Lian, Kewei, et al.
Published: (2025)
One-Step Diffusion for Detail-Rich and Temporally Consistent Video Super-Resolution
by: Sun, Yujing, et al.
Published: (2025)
by: Sun, Yujing, et al.
Published: (2025)
UVCG: Leveraging Temporal Consistency for Universal Video Protection
by: Li, KaiZhou, et al.
Published: (2024)
by: Li, KaiZhou, et al.
Published: (2024)
PISCO: Precise Video Instance Insertion with Sparse Control
by: Gao, Xiangbo, et al.
Published: (2026)
by: Gao, Xiangbo, et al.
Published: (2026)
Pack and Force Your Memory: Long-form and Consistent Video Generation
by: Wu, Xiaofei, et al.
Published: (2025)
by: Wu, Xiaofei, et al.
Published: (2025)
"PhyWorldBench": A Comprehensive Evaluation of Physical Realism in Text-to-Video Models
by: Gu, Jing, et al.
Published: (2025)
by: Gu, Jing, et al.
Published: (2025)
GSV3D: Gaussian Splatting-based Geometric Distillation with Stable Video Diffusion for Single-Image 3D Object Generation
by: Tao, Ye, et al.
Published: (2025)
by: Tao, Ye, et al.
Published: (2025)
VideoCogQA: A Controllable Benchmark for Evaluating Cognitive Abilities in Video-Language Models
by: Li, Chenglin, et al.
Published: (2024)
by: Li, Chenglin, et al.
Published: (2024)
Pre-Trained Video Generative Models as World Simulators
by: He, Haoran, et al.
Published: (2025)
by: He, Haoran, et al.
Published: (2025)
Depth-Consistent 3D Gaussian Splatting via Physical Defocus Modeling and Multi-View Geometric Supervision
by: Deng, Yu, et al.
Published: (2025)
by: Deng, Yu, et al.
Published: (2025)
Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward
by: Zhang, Ruohong, et al.
Published: (2024)
by: Zhang, Ruohong, et al.
Published: (2024)
Video Depth Anything: Consistent Depth Estimation for Super-Long Videos
by: Chen, Sili, et al.
Published: (2025)
by: Chen, Sili, et al.
Published: (2025)
MMWorld: Towards Multi-discipline Multi-faceted World Model Evaluation in Videos
by: He, Xuehai, et al.
Published: (2024)
by: He, Xuehai, et al.
Published: (2024)
Interpreting Physics in Video World Models
by: Joseph, Sonia, et al.
Published: (2026)
by: Joseph, Sonia, et al.
Published: (2026)
Tangram: Benchmark for Evaluating Geometric Element Recognition in Large Multimodal Models
by: Zhang, Chao, et al.
Published: (2024)
by: Zhang, Chao, et al.
Published: (2024)
MoCA-Video: Motion-Aware Concept Alignment for Consistent Video Editing
by: Zhang, Tong, et al.
Published: (2025)
by: Zhang, Tong, et al.
Published: (2025)
Video Flow as Time Series: Discovering Temporal Consistency and Variability for VideoQA
by: Song, Zijie, et al.
Published: (2025)
by: Song, Zijie, et al.
Published: (2025)
DropletVideo: A Dataset and Approach to Explore Integral Spatio-Temporal Consistent Video Generation
by: Zhang, Runze, et al.
Published: (2025)
by: Zhang, Runze, et al.
Published: (2025)
Pandora: Towards General World Model with Natural Language Actions and Video States
by: Xiang, Jiannan, et al.
Published: (2024)
by: Xiang, Jiannan, et al.
Published: (2024)
Similar Items
-
PanoWorld: Geometry-Consistent Panoramic Video World Modeling
by: Jiang, Le, et al.
Published: (2026) -
Geometry Forcing: Marrying Video Diffusion and 3D Representation for Consistent World Modeling
by: Wu, Haoyu, et al.
Published: (2025) -
ShareVerse: Multi-Agent Consistent Video Generation for Shared World Modeling
by: Zhu, Jiayi, et al.
Published: (2026) -
WorldReel: 4D Video Generation with Consistent Geometry and Motion Modeling
by: Fang, Shaoheng, et al.
Published: (2025) -
RelightVid: Temporal-Consistent Diffusion Model for Video Relighting
by: Fang, Ye, et al.
Published: (2025)