PhyGround: Benchmarking Physical Reasoning in Generative World Models
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , , , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866910209471938560 |
|---|---|
| author | Lin, Juyi Akbari, Arash He, Yumei Zhao, Lin Zhang, Haichao Akbari, Arman Xu, Xingchen Lu, Zoe Y. Nan, Enfu Deng, Hokin Yeh, Edmund Ostadabbas, Sarah Fu, Yun Dy, Jennifer Zhao, Pu Wang, Yanzhi |
| author_facet | Lin, Juyi Akbari, Arash He, Yumei Zhao, Lin Zhang, Haichao Akbari, Arman Xu, Xingchen Lu, Zoe Y. Nan, Enfu Deng, Hokin Yeh, Edmund Ostadabbas, Sarah Fu, Yun Dy, Jennifer Zhao, Pu Wang, Yanzhi |
| contents | Generative world models are increasingly used for video generation, where learned simulators are expected to capture the physical rules that govern real-world dynamics. However, evaluating whether generated videos actually follow these rules remains challenging. Existing physics-focused video benchmarks have made important progress, but they still face three key challenges, including the coarse evaluation frameworks that hide law-specific failures, response biases and fatigue that undermine the validity of annotation judgments, and automated evaluators that are insufficiently physics-aware or difficult to audit. To address those challenges, we introduce PhyGround, a criteria-grounded benchmark for evaluating physical reasoning in video generation. The benchmark contains 250 curated prompts, each augmented with an expected physical outcome, and a taxonomy of 13 physical laws across solid-body mechanics, fluid dynamics, and optics. Each law is operationalized through observable sub-questions to enable per-law diagnostics. We evaluate eight modern video generation models through a large-scale, quality-controlled human study, grounded on social science lab experiment design. A total of 459 annotators provided 5,796 complete annotations and over 37.4K fine-grained labels; after quality control, the retained annotations exhibited high split-half model-ranking correlations (Spearman's rho > 0.90). To support reproducible automated evaluation, we release PhyJudge-9B, an open physics-specialized VLM judge. PhyJudge-9B achieves substantially lower aggregate relative bias than Gemini-3.1-Pro (3.3% vs. 16.6%). We release prompts, human annotations, model checkpoints, and evaluation code on the project page https://phyground.github.io/. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2605_10806 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | PhyGround: Benchmarking Physical Reasoning in Generative World Models Lin, Juyi Akbari, Arash He, Yumei Zhao, Lin Zhang, Haichao Akbari, Arman Xu, Xingchen Lu, Zoe Y. Nan, Enfu Deng, Hokin Yeh, Edmund Ostadabbas, Sarah Fu, Yun Dy, Jennifer Zhao, Pu Wang, Yanzhi Computer Vision and Pattern Recognition Artificial Intelligence Machine Learning Generative world models are increasingly used for video generation, where learned simulators are expected to capture the physical rules that govern real-world dynamics. However, evaluating whether generated videos actually follow these rules remains challenging. Existing physics-focused video benchmarks have made important progress, but they still face three key challenges, including the coarse evaluation frameworks that hide law-specific failures, response biases and fatigue that undermine the validity of annotation judgments, and automated evaluators that are insufficiently physics-aware or difficult to audit. To address those challenges, we introduce PhyGround, a criteria-grounded benchmark for evaluating physical reasoning in video generation. The benchmark contains 250 curated prompts, each augmented with an expected physical outcome, and a taxonomy of 13 physical laws across solid-body mechanics, fluid dynamics, and optics. Each law is operationalized through observable sub-questions to enable per-law diagnostics. We evaluate eight modern video generation models through a large-scale, quality-controlled human study, grounded on social science lab experiment design. A total of 459 annotators provided 5,796 complete annotations and over 37.4K fine-grained labels; after quality control, the retained annotations exhibited high split-half model-ranking correlations (Spearman's rho > 0.90). To support reproducible automated evaluation, we release PhyJudge-9B, an open physics-specialized VLM judge. PhyJudge-9B achieves substantially lower aggregate relative bias than Gemini-3.1-Pro (3.3% vs. 16.6%). We release prompts, human annotations, model checkpoints, and evaluation code on the project page https://phyground.github.io/. |
| title | PhyGround: Benchmarking Physical Reasoning in Generative World Models |
| topic | Computer Vision and Pattern Recognition Artificial Intelligence Machine Learning |
| url | https://arxiv.org/abs/2605.10806 |