ORV: 4D Occupancy-centric Robot Video Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Xiuyu, Li, Bohan, Xu, Shaocong, Wang, Nan, Ye, Chongjie, Chen, Zhaoxi, Qin, Minghan, Ding, Yikang, Zhu, Zheng, Jin, Xin, Zhao, Hang, Zhao, Hao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
One View, Many Worlds: Single-Image to 3D Object Meets Generative Domain Randomization for One-Shot 6D Pose Estimation
by: Geng, Zheng, et al.
Published: (2025)
by: Geng, Zheng, et al.
Published: (2025)
Light-X: Generative 4D Video Rendering with Camera and Illumination Control
by: Liu, Tianqi, et al.
Published: (2025)
by: Liu, Tianqi, et al.
Published: (2025)
UniScene: Unified Occupancy-centric Driving Scene Generation
by: Li, Bohan, et al.
Published: (2024)
by: Li, Bohan, et al.
Published: (2024)
Relit-LiVE: Relight Video by Jointly Learning Environment Video
by: Xiao, Weiqing, et al.
Published: (2026)
by: Xiao, Weiqing, et al.
Published: (2026)
TrackOcc: Camera-based 4D Panoptic Occupancy Tracking
by: Chen, Zhuoguang, et al.
Published: (2025)
by: Chen, Zhuoguang, et al.
Published: (2025)
Diffusion-Based Generative Models for 3D Occupancy Prediction in Autonomous Driving
by: Wang, Yunshen, et al.
Published: (2025)
by: Wang, Yunshen, et al.
Published: (2025)
Unifying Appearance Codes and Bilateral Grids for Driving Scene Gaussian Splatting
by: Wang, Nan, et al.
Published: (2025)
by: Wang, Nan, et al.
Published: (2025)
Diffusion Knows Transparency: Repurposing Video Diffusion for Transparent Object Depth and Normal Estimation
by: Xu, Shaocong, et al.
Published: (2025)
by: Xu, Shaocong, et al.
Published: (2025)
GS-Occ3D: Scaling Vision-only Occupancy Reconstruction with Gaussian Splatting
by: Ye, Baijun, et al.
Published: (2025)
by: Ye, Baijun, et al.
Published: (2025)
ArtHOI: Articulated Human-Object Interaction Synthesis by 4D Reconstruction from Video Priors
by: Huang, Zihao, et al.
Published: (2026)
by: Huang, Zihao, et al.
Published: (2026)
Light of Normals: Unified Feature Representation for Universal Photometric Stereo
by: Chen, Houyuan, et al.
Published: (2025)
by: Chen, Houyuan, et al.
Published: (2025)
UniVidX: A Unified Multimodal Framework for Versatile Video Generation via Diffusion Priors
by: Chen, Houyuan, et al.
Published: (2026)
by: Chen, Houyuan, et al.
Published: (2026)
Object-centric 3D Motion Field for Robot Learning from Human Videos
by: Yin, Zhao-Heng, et al.
Published: (2025)
by: Yin, Zhao-Heng, et al.
Published: (2025)
CVT-Occ: Cost Volume Temporal Fusion for 3D Occupancy Prediction
by: Ye, Zhangchen, et al.
Published: (2024)
by: Ye, Zhangchen, et al.
Published: (2024)
ManipDreamer3D : Synthesizing Plausible Robotic Manipulation Video with Occupancy-aware 3D Trajectory
by: Li, Ying, et al.
Published: (2025)
by: Li, Ying, et al.
Published: (2025)
DiST-4D: Disentangled Spatiotemporal Diffusion with Metric Depth for 4D Driving Scene Generation
by: Guo, Jiazhe, et al.
Published: (2025)
by: Guo, Jiazhe, et al.
Published: (2025)
LiDAR-based 4D Occupancy Completion and Forecasting
by: Liu, Xinhao, et al.
Published: (2023)
by: Liu, Xinhao, et al.
Published: (2023)
Occupancy-SLAM: Simultaneously Optimizing Robot Poses and Continuous Occupancy Map
by: Zhao, Liang, et al.
Published: (2024)
by: Zhao, Liang, et al.
Published: (2024)
Humanoid Occupancy: Enabling A Generalized Multimodal Occupancy Perception System on Humanoid Robots
by: Cui, Wei, et al.
Published: (2025)
by: Cui, Wei, et al.
Published: (2025)
Occupancy-SLAM: An Efficient and Robust Algorithm for Simultaneously Optimizing Robot Poses and Occupancy Map
by: Wang, Yingyu, et al.
Published: (2025)
by: Wang, Yingyu, et al.
Published: (2025)
Ego3DT: Tracking Every 3D Object in Ego-centric Videos
by: Hao, Shengyu, et al.
Published: (2024)
by: Hao, Shengyu, et al.
Published: (2024)
Occupancy World Model for Robots
by: Zhang, Zhang, et al.
Published: (2025)
by: Zhang, Zhang, et al.
Published: (2025)
OccVLA: Vision-Language-Action Model with Implicit 3D Occupancy Supervision
by: Liu, Ruixun, et al.
Published: (2025)
by: Liu, Ruixun, et al.
Published: (2025)
NeAR: Coupled Neural Asset-Renderer Stack
by: Li, Hong, et al.
Published: (2025)
by: Li, Hong, et al.
Published: (2025)
Hierarchical Context Alignment with Disentangled Geometric and Temporal Modeling for Semantic Occupancy Prediction
by: Li, Bohan, et al.
Published: (2024)
by: Li, Bohan, et al.
Published: (2024)
Panoramic Multimodal Semantic Occupancy Prediction for Quadruped Robots
by: Zhao, Guoqiang, et al.
Published: (2026)
by: Zhao, Guoqiang, et al.
Published: (2026)
Super4DR: 4D Radar-centric Self-supervised Odometry and Gaussian-based Map Optimization
by: Li, Zhiheng, et al.
Published: (2025)
by: Li, Zhiheng, et al.
Published: (2025)
RadarOcc: Robust 3D Occupancy Prediction with 4D Imaging Radar
by: Ding, Fangqiang, et al.
Published: (2024)
by: Ding, Fangqiang, et al.
Published: (2024)
MonoOcc: Digging into Monocular Semantic Occupancy Prediction
by: Zheng, Yupeng, et al.
Published: (2024)
by: Zheng, Yupeng, et al.
Published: (2024)
Physics-informed Neural Network Predictive Control for Quadruped Locomotion
by: Li, Haolin, et al.
Published: (2025)
by: Li, Haolin, et al.
Published: (2025)
Rethinking Image-to-Video Adaptation: An Object-centric Perspective
by: Qian, Rui, et al.
Published: (2024)
by: Qian, Rui, et al.
Published: (2024)
GenieDrive: Towards Physics-Aware Driving World Model with 4D Occupancy Guided Video Generation
by: Yang, Zhenya, et al.
Published: (2025)
by: Yang, Zhenya, et al.
Published: (2025)
OccupancyDETR: Using DETR for Mixed Dense-sparse 3D Occupancy Prediction
by: Jia, Yupeng, et al.
Published: (2023)
by: Jia, Yupeng, et al.
Published: (2023)
A Unified Framework for Human-centric Point Cloud Video Understanding
by: Xu, Yiteng, et al.
Published: (2024)
by: Xu, Yiteng, et al.
Published: (2024)
LangSurf: Language-Embedded Surface Gaussians for 3D Scene Understanding
by: Li, Hao, et al.
Published: (2024)
by: Li, Hao, et al.
Published: (2024)
OccScene: Semantic Occupancy-based Cross-task Mutual Learning for 3D Scene Generation
by: Li, Bohan, et al.
Published: (2024)
by: Li, Bohan, et al.
Published: (2024)
Hi3DGen: High-fidelity 3D Geometry Generation from Images via Normal Bridging
by: Ye, Chongjie, et al.
Published: (2025)
by: Ye, Chongjie, et al.
Published: (2025)
LoFA: Learning to Predict Personalized Priors for Fast Adaptation of Visual Generative Models
by: Hao, Yiming, et al.
Published: (2025)
by: Hao, Yiming, et al.
Published: (2025)
PAM: A Pose-Appearance-Motion Engine for Sim-to-Real HOI Video Generation
by: Gao, Mingju, et al.
Published: (2026)
by: Gao, Mingju, et al.
Published: (2026)
Track4World: Feedforward World-centric Dense 3D Tracking of All Pixels
by: Lu, Jiahao, et al.
Published: (2026)
by: Lu, Jiahao, et al.
Published: (2026)
Similar Items
-
One View, Many Worlds: Single-Image to 3D Object Meets Generative Domain Randomization for One-Shot 6D Pose Estimation
by: Geng, Zheng, et al.
Published: (2025) -
Light-X: Generative 4D Video Rendering with Camera and Illumination Control
by: Liu, Tianqi, et al.
Published: (2025) -
UniScene: Unified Occupancy-centric Driving Scene Generation
by: Li, Bohan, et al.
Published: (2024) -
Relit-LiVE: Relight Video by Jointly Learning Environment Video
by: Xiao, Weiqing, et al.
Published: (2026) -
TrackOcc: Camera-based 4D Panoptic Occupancy Tracking
by: Chen, Zhuoguang, et al.
Published: (2025)