PhyGround: Benchmarking Physical Reasoning in Generative World Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Lin, Juyi, Akbari, Arash, He, Yumei, Zhao, Lin, Zhang, Haichao, Akbari, Arman, Xu, Xingchen, Lu, Zoe Y., Nan, Enfu, Deng, Hokin, Yeh, Edmund, Ostadabbas, Sarah, Fu, Yun, Dy, Jennifer, Zhao, Pu, Wang, Yanzhi
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910209471938560
author Lin, Juyi
Akbari, Arash
He, Yumei
Zhao, Lin
Zhang, Haichao
Akbari, Arman
Xu, Xingchen
Lu, Zoe Y.
Nan, Enfu
Deng, Hokin
Yeh, Edmund
Ostadabbas, Sarah
Fu, Yun
Dy, Jennifer
Zhao, Pu
Wang, Yanzhi
author_facet Lin, Juyi
Akbari, Arash
He, Yumei
Zhao, Lin
Zhang, Haichao
Akbari, Arman
Xu, Xingchen
Lu, Zoe Y.
Nan, Enfu
Deng, Hokin
Yeh, Edmund
Ostadabbas, Sarah
Fu, Yun
Dy, Jennifer
Zhao, Pu
Wang, Yanzhi
contents Generative world models are increasingly used for video generation, where learned simulators are expected to capture the physical rules that govern real-world dynamics. However, evaluating whether generated videos actually follow these rules remains challenging. Existing physics-focused video benchmarks have made important progress, but they still face three key challenges, including the coarse evaluation frameworks that hide law-specific failures, response biases and fatigue that undermine the validity of annotation judgments, and automated evaluators that are insufficiently physics-aware or difficult to audit. To address those challenges, we introduce PhyGround, a criteria-grounded benchmark for evaluating physical reasoning in video generation. The benchmark contains 250 curated prompts, each augmented with an expected physical outcome, and a taxonomy of 13 physical laws across solid-body mechanics, fluid dynamics, and optics. Each law is operationalized through observable sub-questions to enable per-law diagnostics. We evaluate eight modern video generation models through a large-scale, quality-controlled human study, grounded on social science lab experiment design. A total of 459 annotators provided 5,796 complete annotations and over 37.4K fine-grained labels; after quality control, the retained annotations exhibited high split-half model-ranking correlations (Spearman's rho > 0.90). To support reproducible automated evaluation, we release PhyJudge-9B, an open physics-specialized VLM judge. PhyJudge-9B achieves substantially lower aggregate relative bias than Gemini-3.1-Pro (3.3% vs. 16.6%). We release prompts, human annotations, model checkpoints, and evaluation code on the project page https://phyground.github.io/.
format Preprint
id arxiv_https___arxiv_org_abs_2605_10806
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle PhyGround: Benchmarking Physical Reasoning in Generative World Models
Lin, Juyi
Akbari, Arash
He, Yumei
Zhao, Lin
Zhang, Haichao
Akbari, Arman
Xu, Xingchen
Lu, Zoe Y.
Nan, Enfu
Deng, Hokin
Yeh, Edmund
Ostadabbas, Sarah
Fu, Yun
Dy, Jennifer
Zhao, Pu
Wang, Yanzhi
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Generative world models are increasingly used for video generation, where learned simulators are expected to capture the physical rules that govern real-world dynamics. However, evaluating whether generated videos actually follow these rules remains challenging. Existing physics-focused video benchmarks have made important progress, but they still face three key challenges, including the coarse evaluation frameworks that hide law-specific failures, response biases and fatigue that undermine the validity of annotation judgments, and automated evaluators that are insufficiently physics-aware or difficult to audit. To address those challenges, we introduce PhyGround, a criteria-grounded benchmark for evaluating physical reasoning in video generation. The benchmark contains 250 curated prompts, each augmented with an expected physical outcome, and a taxonomy of 13 physical laws across solid-body mechanics, fluid dynamics, and optics. Each law is operationalized through observable sub-questions to enable per-law diagnostics. We evaluate eight modern video generation models through a large-scale, quality-controlled human study, grounded on social science lab experiment design. A total of 459 annotators provided 5,796 complete annotations and over 37.4K fine-grained labels; after quality control, the retained annotations exhibited high split-half model-ranking correlations (Spearman's rho > 0.90). To support reproducible automated evaluation, we release PhyJudge-9B, an open physics-specialized VLM judge. PhyJudge-9B achieves substantially lower aggregate relative bias than Gemini-3.1-Pro (3.3% vs. 16.6%). We release prompts, human annotations, model checkpoints, and evaluation code on the project page https://phyground.github.io/.
title PhyGround: Benchmarking Physical Reasoning in Generative World Models
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2605.10806