Optimizing Human Pose Estimation Through Focused Human and Joint Regions

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jiao, Yingying, Wang, Zhigang, Liu, Zhenguang, Fan, Shaojing, Wu, Sifan, Wu, Zheqi, Xu, Zhuoyue
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910797988364288
author Jiao, Yingying
Wang, Zhigang
Liu, Zhenguang
Fan, Shaojing
Wu, Sifan
Wu, Zheqi
Xu, Zhuoyue
author_facet Jiao, Yingying
Wang, Zhigang
Liu, Zhenguang
Fan, Shaojing
Wu, Sifan
Wu, Zheqi
Xu, Zhuoyue
contents Human pose estimation has given rise to a broad spectrum of novel and compelling applications, including action recognition, sports analysis, as well as surveillance. However, accurate video pose estimation remains an open challenge. One aspect that has been overlooked so far is that existing methods learn motion clues from all pixels rather than focusing on the target human body, making them easily misled and disrupted by unimportant information such as background changes or movements of other people. Additionally, while the current Transformer-based pose estimation methods has demonstrated impressive performance with global modeling, they struggle with local context perception and precise positional identification. In this paper, we try to tackle these challenges from three aspects: (1) We propose a bilayer Human-Keypoint Mask module that performs coarse-to-fine visual token refinement, which gradually zooms in on the target human body and keypoints while masking out unimportant figure regions. (2) We further introduce a novel deformable cross attention mechanism and a bidirectional separation strategy to adaptively aggregate spatial and temporal motion clues from constrained surrounding contexts. (3) We mathematically formulate the deformable cross attention, constraining that the model focuses solely on the regions centered at the target person body. Empirically, our method achieves state-of-the-art performance on three large-scale benchmark datasets. A remarkable highlight is that our method achieves an 84.8 mean Average Precision (mAP) on the challenging wrist joint, which significantly outperforms the 81.5 mAP achieved by the current state-of-the-art method on the PoseTrack2017 dataset.
format Preprint
id arxiv_https___arxiv_org_abs_2501_14439
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Optimizing Human Pose Estimation Through Focused Human and Joint Regions
Jiao, Yingying
Wang, Zhigang
Liu, Zhenguang
Fan, Shaojing
Wu, Sifan
Wu, Zheqi
Xu, Zhuoyue
Computer Vision and Pattern Recognition
Human pose estimation has given rise to a broad spectrum of novel and compelling applications, including action recognition, sports analysis, as well as surveillance. However, accurate video pose estimation remains an open challenge. One aspect that has been overlooked so far is that existing methods learn motion clues from all pixels rather than focusing on the target human body, making them easily misled and disrupted by unimportant information such as background changes or movements of other people. Additionally, while the current Transformer-based pose estimation methods has demonstrated impressive performance with global modeling, they struggle with local context perception and precise positional identification. In this paper, we try to tackle these challenges from three aspects: (1) We propose a bilayer Human-Keypoint Mask module that performs coarse-to-fine visual token refinement, which gradually zooms in on the target human body and keypoints while masking out unimportant figure regions. (2) We further introduce a novel deformable cross attention mechanism and a bidirectional separation strategy to adaptively aggregate spatial and temporal motion clues from constrained surrounding contexts. (3) We mathematically formulate the deformable cross attention, constraining that the model focuses solely on the regions centered at the target person body. Empirically, our method achieves state-of-the-art performance on three large-scale benchmark datasets. A remarkable highlight is that our method achieves an 84.8 mean Average Precision (mAP) on the challenging wrist joint, which significantly outperforms the 81.5 mAP achieved by the current state-of-the-art method on the PoseTrack2017 dataset.
title Optimizing Human Pose Estimation Through Focused Human and Joint Regions
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2501.14439