Saved in:
| Main Authors: | Wu, Fan, Wei, Jiacheng, Li, Ruibo, Xu, Yi, Li, Junyou, Ye, Deheng, Lin, Guosheng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2512.02793 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PhyCustom: Towards Realistic Physical Customization in Text-to-Image Generation
by: Wu, Fan, et al.
Published: (2025)
by: Wu, Fan, et al.
Published: (2025)
StateSpaceDiffuser: Bringing Long Context to Diffusion World Models
by: Savov, Nedko, et al.
Published: (2025)
by: Savov, Nedko, et al.
Published: (2025)
Weakly and Self-Supervised Class-Agnostic Motion Prediction for Autonomous Driving
by: Li, Ruibo, et al.
Published: (2025)
by: Li, Ruibo, et al.
Published: (2025)
UrbanWorld: An Urban World Model for 3D City Generation
by: Shang, Yu, et al.
Published: (2024)
by: Shang, Yu, et al.
Published: (2024)
Human-in-Context: Unified Cross-Domain 3D Human Motion Modeling via In-Context Learning
by: Liu, Mengyuan, et al.
Published: (2025)
by: Liu, Mengyuan, et al.
Published: (2025)
GigaWorld-Policy: An Efficient Action-Centered World--Action Model
by: Ye, Angen, et al.
Published: (2026)
by: Ye, Angen, et al.
Published: (2026)
Point-In-Context: Understanding Point Cloud via In-Context Learning
by: Liu, Mengyuan, et al.
Published: (2024)
by: Liu, Mengyuan, et al.
Published: (2024)
Puppeteer: Rig and Animate Your 3D Models
by: Song, Chaoyue, et al.
Published: (2025)
by: Song, Chaoyue, et al.
Published: (2025)
Prim2Room: Layout-Controllable Room Mesh Generation from Primitives
by: Feng, Chengzeng, et al.
Published: (2024)
by: Feng, Chengzeng, et al.
Published: (2024)
ShareVerse: Multi-Agent Consistent Video Generation for Shared World Modeling
by: Zhu, Jiayi, et al.
Published: (2026)
by: Zhu, Jiayi, et al.
Published: (2026)
Dense Supervision Propagation for Weakly Supervised Semantic Segmentation on 3D Point Clouds
by: Wei, Jiacheng, et al.
Published: (2021)
by: Wei, Jiacheng, et al.
Published: (2021)
IC-Custom: Diverse Image Customization via In-Context Learning
by: Li, Yaowei, et al.
Published: (2025)
by: Li, Yaowei, et al.
Published: (2025)
GigaWorld-0: World Models as Data Engine to Empower Embodied AI
by: GigaWorld Team, et al.
Published: (2025)
by: GigaWorld Team, et al.
Published: (2025)
DreamDance: Animating Character Art via Inpainting Stable Gaussian Worlds
by: Zhang, Jiaxu, et al.
Published: (2025)
by: Zhang, Jiaxu, et al.
Published: (2025)
DriveDreamer-2: LLM-Enhanced World Models for Diverse Driving Video Generation
by: Zhao, Guosheng, et al.
Published: (2024)
by: Zhao, Guosheng, et al.
Published: (2024)
WonderTurbo: Generating Interactive 3D World in 0.72 Seconds
by: Ni, Chaojun, et al.
Published: (2025)
by: Ni, Chaojun, et al.
Published: (2025)
WorldVLN: Autoregressive World Action Model for Aerial Vision-Language Navigation
by: Zhao, Baining, et al.
Published: (2026)
by: Zhao, Baining, et al.
Published: (2026)
WorldCache: Accelerating World Models for Free via Heterogeneous Token Caching
by: Feng, Weilun, et al.
Published: (2026)
by: Feng, Weilun, et al.
Published: (2026)
DreamWorld: Unified World Modeling in Video Generation
by: Tan, Boming, et al.
Published: (2026)
by: Tan, Boming, et al.
Published: (2026)
WorldGrow: Generating Infinite 3D World
by: Li, Sikuang, et al.
Published: (2025)
by: Li, Sikuang, et al.
Published: (2025)
WorldEval: World Model as Real-World Robot Policies Evaluator
by: Li, Yaxuan, et al.
Published: (2025)
by: Li, Yaxuan, et al.
Published: (2025)
LongScape: Advancing Long-Horizon Embodied World Models with Context-Aware MoE
by: Shang, Yu, et al.
Published: (2025)
by: Shang, Yu, et al.
Published: (2025)
TrackingWorld: World-centric Monocular 3D Tracking of Almost All Pixels
by: Lu, Jiahao, et al.
Published: (2025)
by: Lu, Jiahao, et al.
Published: (2025)
REACTO: Reconstructing Articulated Objects from a Single Video
by: Song, Chaoyue, et al.
Published: (2024)
by: Song, Chaoyue, et al.
Published: (2024)
MathScape: Benchmarking Multimodal Large Language Models in Real-World Mathematical Contexts
by: Liang, Hao, et al.
Published: (2024)
by: Liang, Hao, et al.
Published: (2024)
Light Interaction: Training-Free Inference Acceleration for Interactive Video World Models
by: Lu, Jiacheng, et al.
Published: (2026)
by: Lu, Jiacheng, et al.
Published: (2026)
Latent-WAM: Latent World Action Modeling for End-to-End Autonomous Driving
by: Wang, Linbo, et al.
Published: (2026)
by: Wang, Linbo, et al.
Published: (2026)
Yan: Foundational Interactive Video Generation
by: Ye, Deheng, et al.
Published: (2025)
by: Ye, Deheng, et al.
Published: (2025)
Geometry-Aware Implicit Memory for Video World Models
by: Wei, Zhengxuan, et al.
Published: (2026)
by: Wei, Zhengxuan, et al.
Published: (2026)
HY-World 2.0: A Multi-Modal World Model for Reconstructing, Generating, and Simulating 3D Worlds
by: HY-World, Team, et al.
Published: (2026)
by: HY-World, Team, et al.
Published: (2026)
WildWorld: A Large-Scale Dataset for Dynamic World Modeling with Actions and Explicit State toward Generative ARPG
by: Li, Zhen, et al.
Published: (2026)
by: Li, Zhen, et al.
Published: (2026)
MoDA: Modeling Deformable 3D Objects from Casual Videos
by: Song, Chaoyue, et al.
Published: (2023)
by: Song, Chaoyue, et al.
Published: (2023)
OrientDream: Streamlining Text-to-3D Generation with Explicit Orientation Control
by: Huang, Yuzhong, et al.
Published: (2024)
by: Huang, Yuzhong, et al.
Published: (2024)
Xiaomi Auto World Model: A Joint World Model Integrating Reconstruction and Generation for Autonomous Driving
by: Zhou, Lijun, et al.
Published: (2026)
by: Zhou, Lijun, et al.
Published: (2026)
Sync4D: Video Guided Controllable Dynamics for Physics-Based 4D Generation
by: Fu, Zhoujie, et al.
Published: (2024)
by: Fu, Zhoujie, et al.
Published: (2024)
Generalized Dynamics Generation towards Scannable Physical World Model
by: Li, Yichen, et al.
Published: (2025)
by: Li, Yichen, et al.
Published: (2025)
An Efficient Occupancy World Model via Decoupled Dynamic Flow and Image-assisted Training
by: Zhang, Haiming, et al.
Published: (2024)
by: Zhang, Haiming, et al.
Published: (2024)
GigaBrain-0: A World Model-Powered Vision-Language-Action Model
by: GigaBrain Team, et al.
Published: (2025)
by: GigaBrain Team, et al.
Published: (2025)
OccLE: Label-Efficient 3D Semantic Occupancy Prediction
by: Fang, Naiyu, et al.
Published: (2025)
by: Fang, Naiyu, et al.
Published: (2025)
WorldSimBench: Towards Video Generation Models as World Simulators
by: Qin, Yiran, et al.
Published: (2024)
by: Qin, Yiran, et al.
Published: (2024)
Similar Items
-
PhyCustom: Towards Realistic Physical Customization in Text-to-Image Generation
by: Wu, Fan, et al.
Published: (2025) -
StateSpaceDiffuser: Bringing Long Context to Diffusion World Models
by: Savov, Nedko, et al.
Published: (2025) -
Weakly and Self-Supervised Class-Agnostic Motion Prediction for Autonomous Driving
by: Li, Ruibo, et al.
Published: (2025) -
UrbanWorld: An Urban World Model for 3D City Generation
by: Shang, Yu, et al.
Published: (2024) -
Human-in-Context: Unified Cross-Domain 3D Human Motion Modeling via In-Context Learning
by: Liu, Mengyuan, et al.
Published: (2025)