PhysFlow: Unleashing the Potential of Multi-modal Foundation Models and Video Diffusion for 4D Dynamic Physical Scene Simulation
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Zhuoman, Ye, Weicai, Luximon, Yan, Wan, Pengfei, Zhang, Di |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SIR: Multi-view Inverse Rendering with Decomposable Shadow Under Indoor Intense Lighting
by: Wei, Xiaokang, et al.
Published: (2024)
by: Wei, Xiaokang, et al.
Published: (2024)
PhysFlow: Skin tone transfer for remote heart rate estimation through conditional normalizing flows
by: Comas, Joaquim, et al.
Published: (2024)
by: Comas, Joaquim, et al.
Published: (2024)
FullDiT: Multi-Task Video Generative Foundation Model with Full Attention
by: Ju, Xuan, et al.
Published: (2025)
by: Ju, Xuan, et al.
Published: (2025)
FullDiT2: Efficient In-Context Conditioning for Video Diffusion Transformers
by: He, Xuanhua, et al.
Published: (2025)
by: He, Xuanhua, et al.
Published: (2025)
SketchVideo: Sketch-based Video Generation and Editing
by: Liu, Feng-Lin, et al.
Published: (2025)
by: Liu, Feng-Lin, et al.
Published: (2025)
DifFlow3D: Toward Robust Uncertainty-Aware Scene Flow Estimation with Diffusion Model
by: Liu, Jiuming, et al.
Published: (2023)
by: Liu, Jiuming, et al.
Published: (2023)
Phys4D: Fine-Grained Physics-Consistent 4D Modeling from Video Diffusion
by: Lu, Haoran, et al.
Published: (2026)
by: Lu, Haoran, et al.
Published: (2026)
PhysMaster: Mastering Physical Representation for Video Generation via Reinforcement Learning
by: Ji, Sihui, et al.
Published: (2025)
by: Ji, Sihui, et al.
Published: (2025)
MSGNav: Unleashing the Power of Multi-modal 3D Scene Graph for Zero-Shot Embodied Navigation
by: Huang, Xun, et al.
Published: (2025)
by: Huang, Xun, et al.
Published: (2025)
AdaViewPlanner: Adapting Video Diffusion Models for Viewpoint Planning in 4D Scenes
by: Li, Yu, et al.
Published: (2025)
by: Li, Yu, et al.
Published: (2025)
IPFormer-VideoLLM: Enhancing Multi-modal Video Understanding for Multi-shot Scenes
by: Liang, Yujia, et al.
Published: (2025)
by: Liang, Yujia, et al.
Published: (2025)
Unleashing Network Potentials for Semantic Scene Completion
by: Wang, Fengyun, et al.
Published: (2024)
by: Wang, Fengyun, et al.
Published: (2024)
In-Context Audio Control of Video Diffusion Transformers
by: Liu, Wenze, et al.
Published: (2025)
by: Liu, Wenze, et al.
Published: (2025)
VideoCanvas: Unified Video Completion from Arbitrary Spatiotemporal Patches via In-Context Conditioning
by: Cai, Minghong, et al.
Published: (2025)
by: Cai, Minghong, et al.
Published: (2025)
Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding
by: Yin, Xingyilang, et al.
Published: (2025)
by: Yin, Xingyilang, et al.
Published: (2025)
DreamScene4D: Dynamic Multi-Object Scene Generation from Monocular Videos
by: Chu, Wen-Hsuan, et al.
Published: (2024)
by: Chu, Wen-Hsuan, et al.
Published: (2024)
InterPhys: Physics-aware Human Motion Synthesis in a Dynamic Scene
by: Xing, Chaoyue, et al.
Published: (2026)
by: Xing, Chaoyue, et al.
Published: (2026)
UniVerse: Unleashing the Scene Prior of Video Diffusion Models for Robust Radiance Field Reconstruction
by: Cao, Jin, et al.
Published: (2025)
by: Cao, Jin, et al.
Published: (2025)
SpatialCrafter: Unleashing the Imagination of Video Diffusion Models for Scene Reconstruction from Limited Observations
by: Zhang, Songchun, et al.
Published: (2025)
by: Zhang, Songchun, et al.
Published: (2025)
VINO: A Unified Visual Generator with Interleaved OmniModal Context
by: Chen, Junyi, et al.
Published: (2026)
by: Chen, Junyi, et al.
Published: (2026)
D$^3$FlowSLAM: Self-Supervised Dynamic SLAM with Flow Motion Decomposition and DINO Guidance
by: Yu, Xingyuan, et al.
Published: (2022)
by: Yu, Xingyuan, et al.
Published: (2022)
Any2Caption:Interpreting Any Condition to Caption for Controllable Video Generation
by: Wu, Shengqiong, et al.
Published: (2025)
by: Wu, Shengqiong, et al.
Published: (2025)
H-Flow: Self-supervised Human Scene Flow via Physics-inspired Joint Multi-modal Learning
by: Huang, Zhanbo, et al.
Published: (2026)
by: Huang, Zhanbo, et al.
Published: (2026)
PhysConvex: Physics-Informed 3D Dynamic Convex Radiance Fields for Reconstruction and Simulation
by: Wang, Dan, et al.
Published: (2026)
by: Wang, Dan, et al.
Published: (2026)
VDEGaussian: Video Diffusion Enhanced 4D Gaussian Splatting for Dynamic Urban Scenes Modeling
by: Xiao, Yuru, et al.
Published: (2025)
by: Xiao, Yuru, et al.
Published: (2025)
VMoBA: Mixture-of-Block Attention for Video Diffusion Models
by: Wu, Jianzong, et al.
Published: (2025)
by: Wu, Jianzong, et al.
Published: (2025)
GigaGS: Scaling up Planar-Based 3D Gaussians for Large Scene Surface Reconstruction
by: Chen, Junyi, et al.
Published: (2024)
by: Chen, Junyi, et al.
Published: (2024)
Unleash the Potential of CLIP for Video Highlight Detection
by: Han, Donghoon, et al.
Published: (2024)
by: Han, Donghoon, et al.
Published: (2024)
Phys4DGen: Physics-Compliant 4D Generation with Multi-Material Composition Perception
by: Lin, Jiajing, et al.
Published: (2024)
by: Lin, Jiajing, et al.
Published: (2024)
PhysTalk: Language-driven Real-time Physics in 3D Gaussian Scenes
by: Collorone, Luca, et al.
Published: (2025)
by: Collorone, Luca, et al.
Published: (2025)
CineScene: Implicit 3D as Effective Scene Representation for Cinematic Video Generation
by: Huang, Kaiyi, et al.
Published: (2026)
by: Huang, Kaiyi, et al.
Published: (2026)
PhysChoreo: Physics-Controllable Video Generation with Part-Aware Semantic Grounding
by: Zhang, Haoze, et al.
Published: (2025)
by: Zhang, Haoze, et al.
Published: (2025)
3DPhysVideo: Consistency-Guided Flow SDE for Video Generation via 3D Scene Reconstruction and Physical Simulation
by: Kim, Hwidong, et al.
Published: (2026)
by: Kim, Hwidong, et al.
Published: (2026)
PhysTwin: Physics-Informed Reconstruction and Simulation of Deformable Objects from Videos
by: Jiang, Hanxiao, et al.
Published: (2025)
by: Jiang, Hanxiao, et al.
Published: (2025)
A Reason-then-Describe Instruction Interpreter for Controllable Video Generation
by: Wu, Shengqiong, et al.
Published: (2025)
by: Wu, Shengqiong, et al.
Published: (2025)
PhysGM: Large Physical Gaussian Model for Feed-Forward 4D Synthesis
by: Lv, Chunji, et al.
Published: (2025)
by: Lv, Chunji, et al.
Published: (2025)
Compositional 4D Dynamic Scenes Understanding with Physics Priors for Video Question Answering
by: Wang, Xingrui, et al.
Published: (2024)
by: Wang, Xingrui, et al.
Published: (2024)
DAG: Unleash the Potential of Diffusion Model for Open-Vocabulary 3D Affordance Grounding
by: Wang, Hanqing, et al.
Published: (2025)
by: Wang, Hanqing, et al.
Published: (2025)
Exploring Efficient Foundational Multi-modal Models for Video Summarization
by: Samel, Karan, et al.
Published: (2024)
by: Samel, Karan, et al.
Published: (2024)
Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval
by: Yu, Jiwen, et al.
Published: (2025)
by: Yu, Jiwen, et al.
Published: (2025)
Similar Items
-
SIR: Multi-view Inverse Rendering with Decomposable Shadow Under Indoor Intense Lighting
by: Wei, Xiaokang, et al.
Published: (2024) -
PhysFlow: Skin tone transfer for remote heart rate estimation through conditional normalizing flows
by: Comas, Joaquim, et al.
Published: (2024) -
FullDiT: Multi-Task Video Generative Foundation Model with Full Attention
by: Ju, Xuan, et al.
Published: (2025) -
FullDiT2: Efficient In-Context Conditioning for Video Diffusion Transformers
by: He, Xuanhua, et al.
Published: (2025) -
SketchVideo: Sketch-based Video Generation and Editing
by: Liu, Feng-Lin, et al.
Published: (2025)