ProPhy: Progressive Physical Alignment for Dynamic World Simulation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Zijun, Hu, Panwen, Wang, Jing, Zhang, Terry Jingchen, Cheng, Yuhao, Chen, Long, Yan, Yiqiang, Jiang, Zutao, Li, Hanhui, Liang, Xiaodan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918440262959104
author Wang, Zijun
Hu, Panwen
Wang, Jing
Zhang, Terry Jingchen
Cheng, Yuhao
Chen, Long
Yan, Yiqiang
Jiang, Zutao
Li, Hanhui
Liang, Xiaodan
author_facet Wang, Zijun
Hu, Panwen
Wang, Jing
Zhang, Terry Jingchen
Cheng, Yuhao
Chen, Long
Yan, Yiqiang
Jiang, Zutao
Li, Hanhui
Liang, Xiaodan
contents Recent advances in video generation have shown remarkable potential for constructing world simulators. However, current models still struggle to produce physically consistent results, particularly when handling large-scale or complex dynamics. This limitation arises primarily because existing approaches respond isotropically to physical prompts and neglect the fine-grained alignment between generated content and localized physical cues. To address these challenges, we propose ProPhy, a Progressive Physical Alignment Framework that enables explicit physics-aware conditioning and anisotropic generation. ProPhy employs a two-stage Mixture-of-Physics-Experts mechanism for discriminative physical prior extraction, where Semantic Experts infer semantic-level physical principles from textual descriptions, and Refinement Experts capture token-level physical dynamics. This mechanism allows the model to learn fine-grained, physics-aware video representations that better reflect underlying physical laws. Furthermore, we introduce a physical alignment strategy that transfers the physical reasoning capabilities of vision-language models into the Refinement Experts, facilitating a more accurate representation of dynamic physical phenomena. Extensive experiments on physics-aware video generation benchmarks demonstrate that ProPhy produces more realistic, dynamic, and physically coherent results than existing state-of-the-art methods.
format Preprint
id arxiv_https___arxiv_org_abs_2512_05564
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ProPhy: Progressive Physical Alignment for Dynamic World Simulation
Wang, Zijun
Hu, Panwen
Wang, Jing
Zhang, Terry Jingchen
Cheng, Yuhao
Chen, Long
Yan, Yiqiang
Jiang, Zutao
Li, Hanhui
Liang, Xiaodan
Computer Vision and Pattern Recognition
Recent advances in video generation have shown remarkable potential for constructing world simulators. However, current models still struggle to produce physically consistent results, particularly when handling large-scale or complex dynamics. This limitation arises primarily because existing approaches respond isotropically to physical prompts and neglect the fine-grained alignment between generated content and localized physical cues. To address these challenges, we propose ProPhy, a Progressive Physical Alignment Framework that enables explicit physics-aware conditioning and anisotropic generation. ProPhy employs a two-stage Mixture-of-Physics-Experts mechanism for discriminative physical prior extraction, where Semantic Experts infer semantic-level physical principles from textual descriptions, and Refinement Experts capture token-level physical dynamics. This mechanism allows the model to learn fine-grained, physics-aware video representations that better reflect underlying physical laws. Furthermore, we introduce a physical alignment strategy that transfers the physical reasoning capabilities of vision-language models into the Refinement Experts, facilitating a more accurate representation of dynamic physical phenomena. Extensive experiments on physics-aware video generation benchmarks demonstrate that ProPhy produces more realistic, dynamic, and physically coherent results than existing state-of-the-art methods.
title ProPhy: Progressive Physical Alignment for Dynamic World Simulation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.05564