PyCAT4: A Hierarchical Vision Transformer-based Framework for 3D Human Pose Estimation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Zongyou, Loo, Jonathan, Hou, Yinghan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910259236306944
author Yang, Zongyou
Loo, Jonathan
Hou, Yinghan
author_facet Yang, Zongyou
Loo, Jonathan
Hou, Yinghan
contents Recently, a significant improvement in the accuracy of 3D human pose estimation has been achieved by combining convolutional neural networks (CNNs) with pyramid grid alignment feedback loops. Additionally, innovative breakthroughs have been made in the field of computer vision through the adoption of Transformer-based temporal analysis architectures. Given these advancements, this study aims to deeply optimize and improve the existing Pymaf network architecture. The main innovations of this paper include: (1) Introducing a Transformer feature extraction network layer based on self-attention mechanisms to enhance the capture of low-level features; (2) Enhancing the understanding and capture of temporal signals in video sequences through feature temporal fusion techniques; (3) Implementing spatial pyramid structures to achieve multi-scale feature fusion, effectively balancing feature representations differences across different scales. The new PyCAT4 model obtained in this study is validated through experiments on the COCO and 3DPW datasets. The results demonstrate that the proposed improvement strategies significantly enhance the network's detection capability in human pose estimation, further advancing the development of human pose estimation technology.
format Preprint
id arxiv_https___arxiv_org_abs_2508_02806
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle PyCAT4: A Hierarchical Vision Transformer-based Framework for 3D Human Pose Estimation
Yang, Zongyou
Loo, Jonathan
Hou, Yinghan
Computer Vision and Pattern Recognition
Machine Learning
I.2.10; I.4.8; I.5.4
Recently, a significant improvement in the accuracy of 3D human pose estimation has been achieved by combining convolutional neural networks (CNNs) with pyramid grid alignment feedback loops. Additionally, innovative breakthroughs have been made in the field of computer vision through the adoption of Transformer-based temporal analysis architectures. Given these advancements, this study aims to deeply optimize and improve the existing Pymaf network architecture. The main innovations of this paper include: (1) Introducing a Transformer feature extraction network layer based on self-attention mechanisms to enhance the capture of low-level features; (2) Enhancing the understanding and capture of temporal signals in video sequences through feature temporal fusion techniques; (3) Implementing spatial pyramid structures to achieve multi-scale feature fusion, effectively balancing feature representations differences across different scales. The new PyCAT4 model obtained in this study is validated through experiments on the COCO and 3DPW datasets. The results demonstrate that the proposed improvement strategies significantly enhance the network's detection capability in human pose estimation, further advancing the development of human pose estimation technology.
title PyCAT4: A Hierarchical Vision Transformer-based Framework for 3D Human Pose Estimation
topic Computer Vision and Pattern Recognition
Machine Learning
I.2.10; I.4.8; I.5.4
url https://arxiv.org/abs/2508.02806