DSG-World: Learning a 3D Gaussian World Model from Dual State Videos
Fuente:
arXiv
Saved in:
| Main Authors: | Hu, Wenhao, Wen, Xuexiang, Li, Xi, Wang, Gaoang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
IGFuse: Interactive 3D Gaussian Scene Reconstruction via Multi-Scans Fusion
by: Hu, Wenhao, et al.
Published: (2025)
by: Hu, Wenhao, et al.
Published: (2025)
Pointmap Association and Piecewise-Plane Constraint for Consistent and Compact 3D Gaussian Segmentation Field
by: Hu, Wenhao, et al.
Published: (2025)
by: Hu, Wenhao, et al.
Published: (2025)
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding
by: Xu, Weili, et al.
Published: (2025)
by: Xu, Weili, et al.
Published: (2025)
RecurGS: Interactive Scene Modeling via Discrete-State Recurrent Gaussian Fusion
by: Hu, Wenhao, et al.
Published: (2025)
by: Hu, Wenhao, et al.
Published: (2025)
Hand3R: Online 4D Hand-Scene Reconstruction in the Wild
by: Hu, Wendi, et al.
Published: (2026)
by: Hu, Wendi, et al.
Published: (2026)
Ego3DT: Tracking Every 3D Object in Ego-centric Videos
by: Hao, Shengyu, et al.
Published: (2024)
by: Hao, Shengyu, et al.
Published: (2024)
MovieChat+: Question-aware Sparse Memory for Long Video Question Answering
by: Song, Enxin, et al.
Published: (2024)
by: Song, Enxin, et al.
Published: (2024)
ArtifactWorld: Scaling 3D Gaussian Splatting Artifact Restoration via Video Generation Models
by: Wang, Xinliang, et al.
Published: (2026)
by: Wang, Xinliang, et al.
Published: (2026)
Hi-LSplat: Hierarchical 3D Language Gaussian Splatting
by: Zhan, Chenlu, et al.
Published: (2025)
by: Zhan, Chenlu, et al.
Published: (2025)
GaussianSwap: Animatable Video Face Swapping with 3D Gaussian Splatting
by: Cheng, Xuan, et al.
Published: (2026)
by: Cheng, Xuan, et al.
Published: (2026)
CityCraft: A Real Crafter for 3D City Generation
by: Deng, Jie, et al.
Published: (2024)
by: Deng, Jie, et al.
Published: (2024)
FantasyWorld: Geometry-Consistent World Modeling via Unified Video and 3D Prediction
by: Dai, Yixiang, et al.
Published: (2025)
by: Dai, Yixiang, et al.
Published: (2025)
4DGT: Learning a 4D Gaussian Transformer Using Real-World Monocular Videos
by: Xu, Zhen, et al.
Published: (2025)
by: Xu, Zhen, et al.
Published: (2025)
RDG-GS: Relative Depth Guidance with Gaussian Splatting for Real-time Sparse-View 3D Rendering
by: Zhan, Chenlu, et al.
Published: (2025)
by: Zhan, Chenlu, et al.
Published: (2025)
MPM: A Unified 2D-3D Human Pose Representation via Masked Pose Modeling
by: Zhang, Zhenyu, et al.
Published: (2023)
by: Zhang, Zhenyu, et al.
Published: (2023)
Dual-Branch Graph Transformer Network for 3D Human Mesh Reconstruction from Video
by: Tang, Tao, et al.
Published: (2024)
by: Tang, Tao, et al.
Published: (2024)
FlexWorld: Progressively Expanding 3D Scenes for Flexiable-View Synthesis
by: Chen, Luxi, et al.
Published: (2025)
by: Chen, Luxi, et al.
Published: (2025)
WorldTree: Towards 4D Dynamic Worlds from Monocular Video using Tree-Chains
by: Wang, Qisen, et al.
Published: (2026)
by: Wang, Qisen, et al.
Published: (2026)
World-consistent Video Diffusion with Explicit 3D Modeling
by: Zhang, Qihang, et al.
Published: (2024)
by: Zhang, Qihang, et al.
Published: (2024)
Self-paced Multi-grained Cross-modal Interaction Modeling for Referring Expression Comprehension
by: Miao, Peihan, et al.
Published: (2022)
by: Miao, Peihan, et al.
Published: (2022)
User-Aware Prefix-Tuning is a Good Learner for Personalized Image Captioning
by: Wang, Xuan, et al.
Published: (2023)
by: Wang, Xuan, et al.
Published: (2023)
EgoGaussian: Dynamic Scene Understanding from Egocentric Video with 3D Gaussian Splatting
by: Zhang, Daiwei, et al.
Published: (2024)
by: Zhang, Daiwei, et al.
Published: (2024)
Long-Context State-Space Video World Models
by: Po, Ryan, et al.
Published: (2025)
by: Po, Ryan, et al.
Published: (2025)
Learning from Videos for 3D World: Enhancing MLLMs with 3D Vision Geometry Priors
by: Zheng, Duo, et al.
Published: (2025)
by: Zheng, Duo, et al.
Published: (2025)
Martian World Model: Controllable Video Synthesis with Physically Accurate 3D Reconstructions
by: Li, Longfei, et al.
Published: (2025)
by: Li, Longfei, et al.
Published: (2025)
ContactGaussian-WM: Learning Physics-Grounded World Model from Videos
by: Wang, Meizhong, et al.
Published: (2026)
by: Wang, Meizhong, et al.
Published: (2026)
UrbanWorld: An Urban World Model for 3D City Generation
by: Shang, Yu, et al.
Published: (2024)
by: Shang, Yu, et al.
Published: (2024)
GaussianWorld: Gaussian World Model for Streaming 3D Occupancy Prediction
by: Zuo, Sicheng, et al.
Published: (2024)
by: Zuo, Sicheng, et al.
Published: (2024)
RGBD Objects in the Wild: Scaling Real-World 3D Object Learning from RGB-D Videos
by: Xia, Hongchi, et al.
Published: (2024)
by: Xia, Hongchi, et al.
Published: (2024)
VerseCrafter: Dynamic Realistic Video World Model with 4D Geometric Control
by: Zheng, Sixiao, et al.
Published: (2026)
by: Zheng, Sixiao, et al.
Published: (2026)
Learning Object State Changes in Videos: An Open-World Perspective
by: Xue, Zihui, et al.
Published: (2023)
by: Xue, Zihui, et al.
Published: (2023)
GaussianDWM: 3D Gaussian Driving World Model for Unified Scene Understanding and Multi-Modal Generation
by: Deng, Tianchen, et al.
Published: (2025)
by: Deng, Tianchen, et al.
Published: (2025)
Open3DSG: Open-Vocabulary 3D Scene Graphs from Point Clouds with Queryable Objects and Open-Set Relationships
by: Koch, Sebastian, et al.
Published: (2024)
by: Koch, Sebastian, et al.
Published: (2024)
EA3D: Online Open-World 3D Object Extraction from Streaming Videos
by: Zhou, Xiaoyu, et al.
Published: (2025)
by: Zhou, Xiaoyu, et al.
Published: (2025)
DeepVerse: 4D Autoregressive Video Generation as a World Model
by: Chen, Junyi, et al.
Published: (2025)
by: Chen, Junyi, et al.
Published: (2025)
One World, Dual Timeline: Decoupled Spatio-Temporal Gaussian Scene Graph for 4D Cooperative Driving Reconstruction
by: Chen, Yulong, et al.
Published: (2026)
by: Chen, Yulong, et al.
Published: (2026)
DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT
by: Hu, Xiaotao, et al.
Published: (2024)
by: Hu, Xiaotao, et al.
Published: (2024)
WorldSimBench: Towards Video Generation Models as World Simulators
by: Qin, Yiran, et al.
Published: (2024)
by: Qin, Yiran, et al.
Published: (2024)
MagicWorld: Towards Long-Horizon Stability for Interactive Video World Exploration
by: Li, Guangyuan, et al.
Published: (2025)
by: Li, Guangyuan, et al.
Published: (2025)
Flash Sculptor: Modular 3D Worlds from Objects
by: Hu, Yujia, et al.
Published: (2025)
by: Hu, Yujia, et al.
Published: (2025)
Similar Items
-
IGFuse: Interactive 3D Gaussian Scene Reconstruction via Multi-Scans Fusion
by: Hu, Wenhao, et al.
Published: (2025) -
Pointmap Association and Piecewise-Plane Constraint for Consistent and Compact 3D Gaussian Segmentation Field
by: Hu, Wenhao, et al.
Published: (2025) -
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding
by: Xu, Weili, et al.
Published: (2025) -
RecurGS: Interactive Scene Modeling via Discrete-State Recurrent Gaussian Fusion
by: Hu, Wenhao, et al.
Published: (2025) -
Hand3R: Online 4D Hand-Scene Reconstruction in the Wild
by: Hu, Wendi, et al.
Published: (2026)