PAI-Bench: A Comprehensive Benchmark For Physical AI

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhou, Fengzhe, Huang, Jiannan, Li, Jialuo, Ramanan, Deva, Shi, Humphrey
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912741069946880
author Zhou, Fengzhe
Huang, Jiannan
Li, Jialuo
Ramanan, Deva
Shi, Humphrey
author_facet Zhou, Fengzhe
Huang, Jiannan
Li, Jialuo
Ramanan, Deva
Shi, Humphrey
contents Physical AI aims to develop models that can perceive and predict real-world dynamics; yet, the extent to which current multi-modal large language models and video generative models support these abilities is insufficiently understood. We introduce Physical AI Bench (PAI-Bench), a unified and comprehensive benchmark that evaluates perception and prediction capabilities across video generation, conditional video generation, and video understanding, comprising 2,808 real-world cases with task-aligned metrics designed to capture physical plausibility and domain-specific reasoning. Our study provides a systematic assessment of recent models and shows that video generative models, despite strong visual fidelity, often struggle to maintain physically coherent dynamics, while multi-modal large language models exhibit limited performance in forecasting and causal interpretation. These observations suggest that current systems are still at an early stage in handling the perceptual and predictive demands of Physical AI. In summary, PAI-Bench establishes a realistic foundation for evaluating Physical AI and highlights key gaps that future systems must address.
format Preprint
id arxiv_https___arxiv_org_abs_2512_01989
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle PAI-Bench: A Comprehensive Benchmark For Physical AI
Zhou, Fengzhe
Huang, Jiannan
Li, Jialuo
Ramanan, Deva
Shi, Humphrey
Computer Vision and Pattern Recognition
Physical AI aims to develop models that can perceive and predict real-world dynamics; yet, the extent to which current multi-modal large language models and video generative models support these abilities is insufficiently understood. We introduce Physical AI Bench (PAI-Bench), a unified and comprehensive benchmark that evaluates perception and prediction capabilities across video generation, conditional video generation, and video understanding, comprising 2,808 real-world cases with task-aligned metrics designed to capture physical plausibility and domain-specific reasoning. Our study provides a systematic assessment of recent models and shows that video generative models, despite strong visual fidelity, often struggle to maintain physically coherent dynamics, while multi-modal large language models exhibit limited performance in forecasting and causal interpretation. These observations suggest that current systems are still at an early stage in handling the perceptual and predictive demands of Physical AI. In summary, PAI-Bench establishes a realistic foundation for evaluating Physical AI and highlights key gaps that future systems must address.
title PAI-Bench: A Comprehensive Benchmark For Physical AI
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.01989