Phys4D: Fine-Grained Physics-Consistent 4D Modeling from Video Diffusion
Fuente:
arXiv
Saved in:
| Main Authors: | Lu, Haoran, Wu, Shang, Zhang, Jianshu, Su, Maojiang, Ye, Guo, Xu, Chenwei, Lu, Lie, Maneriker, Pranav, Du, Fan, Li, Manling, Wang, Zhaoran, Liu, Han |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PhyPrompt: RL-based Prompt Refinement for Physically Plausible Text-to-Video Generation
by: Wu, Shang, et al.
Published: (2026)
by: Wu, Shang, et al.
Published: (2026)
Towards Sparse Video Understanding and Reasoning
by: Xu, Chenwei, et al.
Published: (2026)
by: Xu, Chenwei, et al.
Published: (2026)
Generalizable Coarse-to-Fine Robot Manipulation via Language-Aligned 3D Keypoints
by: Hu, Jianshu, et al.
Published: (2025)
by: Hu, Jianshu, et al.
Published: (2025)
ConsisVLA-4D: Advancing Spatiotemporal Consistency in Efficient 3D-Perception and 4D-Reasoning for Robotic Manipulation
by: Li, Wei, et al.
Published: (2026)
by: Li, Wei, et al.
Published: (2026)
IKEA Manuals at Work: 4D Grounding of Assembly Instructions on Internet Videos
by: Liu, Yunong, et al.
Published: (2024)
by: Liu, Yunong, et al.
Published: (2024)
Learning to Feel the Future: DreamTacVLA for Contact-Rich Manipulation
by: Ye, Guo, et al.
Published: (2025)
by: Ye, Guo, et al.
Published: (2025)
ManiCM: Real-time 3D Diffusion Policy via Consistency Model for Robotic Manipulation
by: Lu, Guanxing, et al.
Published: (2024)
by: Lu, Guanxing, et al.
Published: (2024)
3D-CovDiffusion: 3D-Aware Diffusion Policy for Coverage Path Planning
by: Chen, Chenyuan, et al.
Published: (2025)
by: Chen, Chenyuan, et al.
Published: (2025)
MagicTac: A Novel High-Resolution 3D Multi-layer Grid-Based Tactile Sensor
by: Fan, Wen, et al.
Published: (2024)
by: Fan, Wen, et al.
Published: (2024)
3D Vision-tactile Reconstruction from Infrared and Visible Images for Robotic Fine-grained Tactile Perception
by: Lin, Yuankai, et al.
Published: (2025)
by: Lin, Yuankai, et al.
Published: (2025)
Language-Model-Assisted Bi-Level Programming for Reward Learning from Internet Videos
by: Mahesheka, Harsh, et al.
Published: (2024)
by: Mahesheka, Harsh, et al.
Published: (2024)
A Dexterous and Compliant Gripper With Soft Hydraulic Actuation for Microgravity Manipulation
by: Su, William, et al.
Published: (2026)
by: Su, William, et al.
Published: (2026)
VG4D: Vision-Language Model Goes 4D Video Recognition
by: Deng, Zhichao, et al.
Published: (2024)
by: Deng, Zhichao, et al.
Published: (2024)
ELAN4D: Embodiment-Centric 4D Supervision for Vision-Language-Action Models via Plug-and-Play Adaptation
by: He, Zeyuan, et al.
Published: (2026)
by: He, Zeyuan, et al.
Published: (2026)
PhysReaction: Physically Plausible Real-Time Humanoid Reaction Synthesis via Forward Dynamics Guided 4D Imitation
by: Liu, Yunze, et al.
Published: (2024)
by: Liu, Yunze, et al.
Published: (2024)
GEM-4D: Geometry-Enhanced Video World Models for Robot Manipulation
by: Zhou, Kaichen, et al.
Published: (2026)
by: Zhou, Kaichen, et al.
Published: (2026)
Open-Ended Multi-Modal Relational Reasoning for Video Question Answering
by: Luo, Haozheng, et al.
Published: (2020)
by: Luo, Haozheng, et al.
Published: (2020)
DSSP: Diffusion State Space Policy with Full-History Encoding
by: Guan, Zhiyuan, et al.
Published: (2026)
by: Guan, Zhiyuan, et al.
Published: (2026)
ConRFT: A Reinforced Fine-tuning Method for VLA Models via Consistency Policy
by: Chen, Yuhui, et al.
Published: (2025)
by: Chen, Yuhui, et al.
Published: (2025)
ZING-3D: Zero-shot Incremental 3D Scene Graphs via Vision-Language Models
by: Saxena, Pranav, et al.
Published: (2025)
by: Saxena, Pranav, et al.
Published: (2025)
Motion Before Action: Diffusing Object Motion as Manipulation Condition
by: Su, Yue, et al.
Published: (2024)
by: Su, Yue, et al.
Published: (2024)
Kinema4D: Kinematic 4D World Modeling for Spatiotemporal Embodied Simulation
by: Xu, Mutian, et al.
Published: (2026)
by: Xu, Mutian, et al.
Published: (2026)
Globally Consistent RGB-D SLAM with 2D Gaussian Splatting
by: Zhong, Xingguang, et al.
Published: (2025)
by: Zhong, Xingguang, et al.
Published: (2025)
EFEAR-4D: Ego-Velocity Filtering for Efficient and Accurate 4D radar Odometry
by: Wu, Xiaoyi, et al.
Published: (2024)
by: Wu, Xiaoyi, et al.
Published: (2024)
Advancing Object Goal Navigation Through LLM-enhanced Object Affinities Transfer
by: Lin, Mengying, et al.
Published: (2024)
by: Lin, Mengying, et al.
Published: (2024)
DreamAvoid: Critical-Phase Test-Time Dreaming to Avoid Failures in VLA Policies
by: Fan, Xianzhe, et al.
Published: (2026)
by: Fan, Xianzhe, et al.
Published: (2026)
DNOI-4DRO: Deep 4D Radar Odometry with Differentiable Neural-Optimization Iterations
by: Lu, Shouyi, et al.
Published: (2025)
by: Lu, Shouyi, et al.
Published: (2025)
CrystalTac: 3D-Printed Vision-Based Tactile Sensor Family through Rapid Monolithic Manufacturing Technique
by: Fan, Wen, et al.
Published: (2024)
by: Fan, Wen, et al.
Published: (2024)
PhysPose: Refining 6D Object Poses with Physical Constraints
by: Malenický, Martin, et al.
Published: (2025)
by: Malenický, Martin, et al.
Published: (2025)
Diffusion Stabilizer Policy for Automated Surgical Robot Manipulations
by: Ho, Chonlam, et al.
Published: (2025)
by: Ho, Chonlam, et al.
Published: (2025)
Koopman Operator Based Linear Model Predictive Control for 2D Quadruped Trotting, Bounding, and Gait Transition
by: Yang, Chun-Ming, et al.
Published: (2025)
by: Yang, Chun-Ming, et al.
Published: (2025)
Kinematics-Aware Diffusion Policy with Consistent 3D Observation and Action Space for Whole-Arm Robotic Manipulation
by: Lv, Kangchen, et al.
Published: (2025)
by: Lv, Kangchen, et al.
Published: (2025)
A 4D Radar Camera Extrinsic Calibration Tool Based on 3D Uncertainty Perspective N Points
by: Cao, Chuan, et al.
Published: (2025)
by: Cao, Chuan, et al.
Published: (2025)
Hamilton--Jacobi Reachability for Spacecraft Collision Avoidance
by: Hui, Larry, et al.
Published: (2026)
by: Hui, Larry, et al.
Published: (2026)
Learning to Generate 4D LiDAR Sequences
by: Liang, Ao, et al.
Published: (2025)
by: Liang, Ao, et al.
Published: (2025)
Learning Fine-Grained Correspondence with Cross-Perspective Perception for Open-Vocabulary 6D Object Pose Estimation
by: Qin, Yu, et al.
Published: (2026)
by: Qin, Yu, et al.
Published: (2026)
Hierarchical GNN-Based Multi-Agent Learning for Dynamic Queue-Jump Lane and Emergency Vehicle Corridor Formation
by: Su, Haoran
Published: (2026)
by: Su, Haoran
Published: (2026)
D$^2$GSLAM: 4D Dynamic Gaussian Splatting SLAM
by: Zhu, Siting, et al.
Published: (2025)
by: Zhu, Siting, et al.
Published: (2025)
pRRTC: GPU-Parallel RRT-Connect for Fast, Consistent, and Low-Cost Motion Planning
by: Huang, Chih H., et al.
Published: (2025)
by: Huang, Chih H., et al.
Published: (2025)
PhysMem: Scaling Test-Time Memory for Embodied Physical Reasoning
by: Li, Haoyang, et al.
Published: (2026)
by: Li, Haoyang, et al.
Published: (2026)
Similar Items
-
PhyPrompt: RL-based Prompt Refinement for Physically Plausible Text-to-Video Generation
by: Wu, Shang, et al.
Published: (2026) -
Towards Sparse Video Understanding and Reasoning
by: Xu, Chenwei, et al.
Published: (2026) -
Generalizable Coarse-to-Fine Robot Manipulation via Language-Aligned 3D Keypoints
by: Hu, Jianshu, et al.
Published: (2025) -
ConsisVLA-4D: Advancing Spatiotemporal Consistency in Efficient 3D-Perception and 4D-Reasoning for Robotic Manipulation
by: Li, Wei, et al.
Published: (2026) -
IKEA Manuals at Work: 4D Grounding of Assembly Instructions on Internet Videos
by: Liu, Yunong, et al.
Published: (2024)