Physion-Eval: Evaluating Physical Realism in Generated Video via Human Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Qin, Jing, Peiyu, Yu, Hong-Xing, Ding, Fangqiang, Nie, Fan, Wang, Weimin, Du, Yilun, Zou, James, Wu, Jiajun, Shuai, Bing |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
"PhyWorldBench": A Comprehensive Evaluation of Physical Realism in Text-to-Video Models
by: Gu, Jing, et al.
Published: (2025)
by: Gu, Jing, et al.
Published: (2025)
Seeking Physics in Diffusion Noise
by: Tang, Chujun, et al.
Published: (2026)
by: Tang, Chujun, et al.
Published: (2026)
VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos
by: Song, Tingyu, et al.
Published: (2025)
by: Song, Tingyu, et al.
Published: (2025)
Ctrl-VI: Controllable Video Synthesis via Variational Inference
by: Duan, Haoyi, et al.
Published: (2025)
by: Duan, Haoyi, et al.
Published: (2025)
VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation
by: Ma, Wentao, et al.
Published: (2025)
by: Ma, Wentao, et al.
Published: (2025)
HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation
by: Yu, Zhaojian, et al.
Published: (2024)
by: Yu, Zhaojian, et al.
Published: (2024)
milliFlow: Scene Flow Estimation on mmWave Radar Point Cloud for Human Motion Sensing
by: Ding, Fangqiang, et al.
Published: (2023)
by: Ding, Fangqiang, et al.
Published: (2023)
ZeroHSI: Zero-Shot 4D Human-Scene Interaction by Video Generation
by: Li, Hongjie, et al.
Published: (2024)
by: Li, Hongjie, et al.
Published: (2024)
How good nnU-Net for Segmenting Cardiac MRI: A Comprehensive Evaluation
by: Gunawardhana, Malitha, et al.
Published: (2024)
by: Gunawardhana, Malitha, et al.
Published: (2024)
EvoLM: In Search of Lost Language Model Training Dynamics
by: Qi, Zhenting, et al.
Published: (2025)
by: Qi, Zhenting, et al.
Published: (2025)
Generalizable Reasoning through Compositional Energy Minimization
by: Oarga, Alexandru, et al.
Published: (2025)
by: Oarga, Alexandru, et al.
Published: (2025)
SedarEval: Automated Evaluation using Self-Adaptive Rubrics
by: Fan, Zhiyuan, et al.
Published: (2025)
by: Fan, Zhiyuan, et al.
Published: (2025)
DocMath-Eval: Evaluating Math Reasoning Capabilities of LLMs in Understanding Long and Specialized Documents
by: Zhao, Yilun, et al.
Published: (2023)
by: Zhao, Yilun, et al.
Published: (2023)
VideoEval: Comprehensive Benchmark Suite for Low-Cost Evaluation of Video Foundation Model
by: Li, Xinhao, et al.
Published: (2024)
by: Li, Xinhao, et al.
Published: (2024)
VideoGen-Eval: Agent-based System for Video Generation Evaluation
by: Yang, Yuhang, et al.
Published: (2025)
by: Yang, Yuhang, et al.
Published: (2025)
Disentangling Foreground and Background Motion for Enhanced Realism in Human Video Generation
by: Liu, Jinlin, et al.
Published: (2024)
by: Liu, Jinlin, et al.
Published: (2024)
PsychEval: A Multi-Session and Multi-Therapy Benchmark for High-Realism AI Psychological Counselor
by: Pan, Qianjun, et al.
Published: (2026)
by: Pan, Qianjun, et al.
Published: (2026)
Grounding Video Models to Actions through Goal Conditioned Exploration
by: Luo, Yunhao, et al.
Published: (2024)
by: Luo, Yunhao, et al.
Published: (2024)
Reasoning with Sampling: Your Base Model is Smarter Than You Think
by: Karan, Aayush, et al.
Published: (2025)
by: Karan, Aayush, et al.
Published: (2025)
Preserving Source Video Realism: High-Fidelity Face Swapping for Cinematic Quality
by: Luo, Zekai, et al.
Published: (2025)
by: Luo, Zekai, et al.
Published: (2025)
HuM-Eval: A Coarse-to-Fine Framework for Human-Centric Video Evaluation
by: Zhang, Bingzi, et al.
Published: (2026)
by: Zhang, Bingzi, et al.
Published: (2026)
M4Human: A Large-Scale Multimodal mmWave Radar Benchmark for Human Mesh Reconstruction
by: Fan, Junqiao, et al.
Published: (2025)
by: Fan, Junqiao, et al.
Published: (2025)
BodyMetric: Evaluating the Realism of Human Bodies in Text-to-Image Generation
by: Andreou, Nefeli, et al.
Published: (2024)
by: Andreou, Nefeli, et al.
Published: (2024)
The Case against Evaluative Realism
by: Dan LÓPEZ DE SA
Published: (2006)
by: Dan LÓPEZ DE SA
Published: (2006)
Temporal Realism Evaluation of Generated Videos Using Compressed-Domain Motion Vectors
by: Cakiroglu, Mert Onur, et al.
Published: (2025)
by: Cakiroglu, Mert Onur, et al.
Published: (2025)
BotEval: Facilitating Interactive Human Evaluation
by: Cho, Hyundong, et al.
Published: (2024)
by: Cho, Hyundong, et al.
Published: (2024)
RadarGen: Automotive Radar Point Cloud Generation from Cameras
by: Borreda, Tomer, et al.
Published: (2025)
by: Borreda, Tomer, et al.
Published: (2025)
RaTrack: Moving Object Detection and Tracking with 4D Radar Point Cloud
by: Pan, Zhijun, et al.
Published: (2023)
by: Pan, Zhijun, et al.
Published: (2023)
RealWonder: Real-Time Physical Action-Conditioned Video Generation
by: Liu, Wei, et al.
Published: (2026)
by: Liu, Wei, et al.
Published: (2026)
Thinking with Spatial Code for Physical-World Video Reasoning
by: Chen, Jieneng, et al.
Published: (2026)
by: Chen, Jieneng, et al.
Published: (2026)
VideoARM: Agentic Reasoning over Hierarchical Memory for Long-Form Video Understanding
by: Yin, Yufei, et al.
Published: (2025)
by: Yin, Yufei, et al.
Published: (2025)
3DSPA: A 3D Semantic Point Autoencoder for Evaluating Video Realism
by: Chandna, Bhavik, et al.
Published: (2026)
by: Chandna, Bhavik, et al.
Published: (2026)
FlagEval Findings Report: A Preliminary Evaluation of Large Reasoning Models on Automatically Verifiable Textual and Visual Questions
by: Qin, Bowen, et al.
Published: (2025)
by: Qin, Bowen, et al.
Published: (2025)
BatchEval: Towards Human-like Text Evaluation
by: Yuan, Peiwen, et al.
Published: (2023)
by: Yuan, Peiwen, et al.
Published: (2023)
AgentsEval: Clinically Faithful Evaluation of Medical Imaging Reports via Multi-Agent Reasoning
by: Fu, Suzhong, et al.
Published: (2026)
by: Fu, Suzhong, et al.
Published: (2026)
Design of A Low-Latency and Parallelizable SVD Dataflow Architecture on FPGA
by: Du, Fangqiang, et al.
Published: (2025)
by: Du, Fangqiang, et al.
Published: (2025)
An Integrated Statistical‐Physical‐Machine Learning Framework: Quantifying Human‐Induced Terrestrial Water Storage Loss
by: Yifan Huang, et al.
Published: (2025)
by: Yifan Huang, et al.
Published: (2025)
Learning Iterative Reasoning through Energy Diffusion
by: Du, Yilun, et al.
Published: (2024)
by: Du, Yilun, et al.
Published: (2024)
Is Visual Realism Enough? Evaluating Gait Biometric Fidelity in Generative AI Human Animation
by: DeAndres-Tame, Ivan, et al.
Published: (2025)
by: DeAndres-Tame, Ivan, et al.
Published: (2025)
Dark Matter Realism: How Referential Semantics Restricts Realism in Contemporary Fundamental Physics
by: Allzén, Simon
Published: (2025)
by: Allzén, Simon
Published: (2025)
Similar Items
-
"PhyWorldBench": A Comprehensive Evaluation of Physical Realism in Text-to-Video Models
by: Gu, Jing, et al.
Published: (2025) -
Seeking Physics in Diffusion Noise
by: Tang, Chujun, et al.
Published: (2026) -
VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos
by: Song, Tingyu, et al.
Published: (2025) -
Ctrl-VI: Controllable Video Synthesis via Variational Inference
by: Duan, Haoyi, et al.
Published: (2025) -
VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation
by: Ma, Wentao, et al.
Published: (2025)