DeforHMR: Vision Transformer with Deformable Cross-Attention for 3D Human Mesh Recovery

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Heo, Jaewoo, Hu, George, Wang, Zeyu, Yeung-Levy, Serena
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929595032272896
author Heo, Jaewoo
Hu, George
Wang, Zeyu
Yeung-Levy, Serena
author_facet Heo, Jaewoo
Hu, George
Wang, Zeyu
Yeung-Levy, Serena
contents Human Mesh Recovery (HMR) is an important yet challenging problem with applications across various domains including motion capture, augmented reality, and biomechanics. Accurately predicting human pose parameters from a single image remains a challenging 3D computer vision task. In this work, we introduce DeforHMR, a novel regression-based monocular HMR framework designed to enhance the prediction of human pose parameters using deformable attention transformers. DeforHMR leverages a novel query-agnostic deformable cross-attention mechanism within the transformer decoder to effectively regress the visual features extracted from a frozen pretrained vision transformer (ViT) encoder. The proposed deformable cross-attention mechanism allows the model to attend to relevant spatial features more flexibly and in a data-dependent manner. Equipped with a transformer decoder capable of spatially-nuanced attention, DeforHMR achieves state-of-the-art performance for single-frame regression-based methods on the widely used 3D HMR benchmarks 3DPW and RICH. By pushing the boundary on the field of 3D human mesh recovery through deformable attention, we introduce an new, effective paradigm for decoding local spatial information from large pretrained vision encoders in computer vision.
format Preprint
id arxiv_https___arxiv_org_abs_2411_11214
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle DeforHMR: Vision Transformer with Deformable Cross-Attention for 3D Human Mesh Recovery
Heo, Jaewoo
Hu, George
Wang, Zeyu
Yeung-Levy, Serena
Computer Vision and Pattern Recognition
Human Mesh Recovery (HMR) is an important yet challenging problem with applications across various domains including motion capture, augmented reality, and biomechanics. Accurately predicting human pose parameters from a single image remains a challenging 3D computer vision task. In this work, we introduce DeforHMR, a novel regression-based monocular HMR framework designed to enhance the prediction of human pose parameters using deformable attention transformers. DeforHMR leverages a novel query-agnostic deformable cross-attention mechanism within the transformer decoder to effectively regress the visual features extracted from a frozen pretrained vision transformer (ViT) encoder. The proposed deformable cross-attention mechanism allows the model to attend to relevant spatial features more flexibly and in a data-dependent manner. Equipped with a transformer decoder capable of spatially-nuanced attention, DeforHMR achieves state-of-the-art performance for single-frame regression-based methods on the widely used 3D HMR benchmarks 3DPW and RICH. By pushing the boundary on the field of 3D human mesh recovery through deformable attention, we introduce an new, effective paradigm for decoding local spatial information from large pretrained vision encoders in computer vision.
title DeforHMR: Vision Transformer with Deformable Cross-Attention for 3D Human Mesh Recovery
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2411.11214