Refined Temporal Pyramidal Compression-and-Amplification Transformer for 3D Human Pose Estimation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Hanbing, Xiang, Wangmeng, He, Jun-Yan, Cheng, Zhi-Qi, Luo, Bin, Geng, Yifeng, Xie, Xuansong
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911770608664576
author Liu, Hanbing
Xiang, Wangmeng
He, Jun-Yan
Cheng, Zhi-Qi
Luo, Bin
Geng, Yifeng
Xie, Xuansong
author_facet Liu, Hanbing
Xiang, Wangmeng
He, Jun-Yan
Cheng, Zhi-Qi
Luo, Bin
Geng, Yifeng
Xie, Xuansong
contents Accurately estimating the 3D pose of humans in video sequences requires both accuracy and a well-structured architecture. With the success of transformers, we introduce the Refined Temporal Pyramidal Compression-and-Amplification (RTPCA) transformer. Exploiting the temporal dimension, RTPCA extends intra-block temporal modeling via its Temporal Pyramidal Compression-and-Amplification (TPCA) structure and refines inter-block feature interaction with a Cross-Layer Refinement (XLR) module. In particular, TPCA block exploits a temporal pyramid paradigm, reinforcing key and value representation capabilities and seamlessly extracting spatial semantics from motion sequences. We stitch these TPCA blocks with XLR that promotes rich semantic representation through continuous interaction of queries, keys, and values. This strategy embodies early-stage information with current flows, addressing typical deficits in detail and stability seen in other transformer-based methods. We demonstrate the effectiveness of RTPCA by achieving state-of-the-art results on Human3.6M, HumanEva-I, and MPI-INF-3DHP benchmarks with minimal computational overhead. The source code is available at https://github.com/hbing-l/RTPCA.
format Preprint
id arxiv_https___arxiv_org_abs_2309_01365
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Refined Temporal Pyramidal Compression-and-Amplification Transformer for 3D Human Pose Estimation
Liu, Hanbing
Xiang, Wangmeng
He, Jun-Yan
Cheng, Zhi-Qi
Luo, Bin
Geng, Yifeng
Xie, Xuansong
Computer Vision and Pattern Recognition
Artificial Intelligence
Accurately estimating the 3D pose of humans in video sequences requires both accuracy and a well-structured architecture. With the success of transformers, we introduce the Refined Temporal Pyramidal Compression-and-Amplification (RTPCA) transformer. Exploiting the temporal dimension, RTPCA extends intra-block temporal modeling via its Temporal Pyramidal Compression-and-Amplification (TPCA) structure and refines inter-block feature interaction with a Cross-Layer Refinement (XLR) module. In particular, TPCA block exploits a temporal pyramid paradigm, reinforcing key and value representation capabilities and seamlessly extracting spatial semantics from motion sequences. We stitch these TPCA blocks with XLR that promotes rich semantic representation through continuous interaction of queries, keys, and values. This strategy embodies early-stage information with current flows, addressing typical deficits in detail and stability seen in other transformer-based methods. We demonstrate the effectiveness of RTPCA by achieving state-of-the-art results on Human3.6M, HumanEva-I, and MPI-INF-3DHP benchmarks with minimal computational overhead. The source code is available at https://github.com/hbing-l/RTPCA.
title Refined Temporal Pyramidal Compression-and-Amplification Transformer for 3D Human Pose Estimation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2309.01365