PhysPT: Physics-aware Pretrained Transformer for Estimating Human Dynamics from Monocular Videos

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Zhang, Yufei, Kephart, Jeffrey O., Cui, Zijun, Ji, Qiang
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866916196205461504
author Zhang, Yufei
Kephart, Jeffrey O.
Cui, Zijun
Ji, Qiang
author_facet Zhang, Yufei
Kephart, Jeffrey O.
Cui, Zijun
Ji, Qiang
contents While current methods have shown promising progress on estimating 3D human motion from monocular videos, their motion estimates are often physically unrealistic because they mainly consider kinematics. In this paper, we introduce Physics-aware Pretrained Transformer (PhysPT), which improves kinematics-based motion estimates and infers motion forces. PhysPT exploits a Transformer encoder-decoder backbone to effectively learn human dynamics in a self-supervised manner. Moreover, it incorporates physics principles governing human motion. Specifically, we build a physics-based body representation and contact force model. We leverage them to impose novel physics-inspired training losses (i.e., force loss, contact loss, and Euler-Lagrange loss), enabling PhysPT to capture physical properties of the human body and the forces it experiences. Experiments demonstrate that, once trained, PhysPT can be directly applied to kinematics-based estimates to significantly enhance their physical plausibility and generate favourable motion forces. Furthermore, we show that these physically meaningful quantities translate into improved accuracy of an important downstream task: human action recognition.
format Preprint
id arxiv_https___arxiv_org_abs_2404_04430
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle PhysPT: Physics-aware Pretrained Transformer for Estimating Human Dynamics from Monocular Videos
Zhang, Yufei
Kephart, Jeffrey O.
Cui, Zijun
Ji, Qiang
Computer Vision and Pattern Recognition
While current methods have shown promising progress on estimating 3D human motion from monocular videos, their motion estimates are often physically unrealistic because they mainly consider kinematics. In this paper, we introduce Physics-aware Pretrained Transformer (PhysPT), which improves kinematics-based motion estimates and infers motion forces. PhysPT exploits a Transformer encoder-decoder backbone to effectively learn human dynamics in a self-supervised manner. Moreover, it incorporates physics principles governing human motion. Specifically, we build a physics-based body representation and contact force model. We leverage them to impose novel physics-inspired training losses (i.e., force loss, contact loss, and Euler-Lagrange loss), enabling PhysPT to capture physical properties of the human body and the forces it experiences. Experiments demonstrate that, once trained, PhysPT can be directly applied to kinematics-based estimates to significantly enhance their physical plausibility and generate favourable motion forces. Furthermore, we show that these physically meaningful quantities translate into improved accuracy of an important downstream task: human action recognition.
title PhysPT: Physics-aware Pretrained Transformer for Estimating Human Dynamics from Monocular Videos
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2404.04430