Efficient Diffusion-Based 3D Human Pose Estimation with Hierarchical Temporal Pruning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Bi, Yuquan, Wang, Hongsong, Shi, Xinli, Gui, Zhipeng, Gui, Jie, Tang, Yuan Yan
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917320077606912
author Bi, Yuquan
Wang, Hongsong
Shi, Xinli
Gui, Zhipeng
Gui, Jie
Tang, Yuan Yan
author_facet Bi, Yuquan
Wang, Hongsong
Shi, Xinli
Gui, Zhipeng
Gui, Jie
Tang, Yuan Yan
contents Diffusion models have demonstrated strong capabilities in generating high-fidelity 3D human poses, yet their iterative nature and multi-hypothesis requirements incur substantial computational cost. In this paper, we propose an Efficient Diffusion-Based 3D Human Pose Estimation framework with a Hierarchical Temporal Pruning (HTP) strategy, which dynamically prunes redundant pose tokens across both frame and semantic levels while preserving critical motion dynamics. HTP operates in a staged, top-down manner: (1) Temporal Correlation-Enhanced Pruning (TCEP) identifies essential frames by analyzing inter-frame motion correlations through adaptive temporal graph construction; (2) Sparse-Focused Temporal MHSA (SFT MHSA) leverages the resulting frame-level sparsity to reduce attention computation, focusing on motion-relevant tokens; and (3) Mask-Guided Pose Token Pruner (MGPTP) performs fine-grained semantic pruning via clustering, retaining only the most informative pose tokens. Experiments on Human3.6M and MPI-INF-3DHP show that HTP reduces training MACs by 38.5\%, inference MACs by 56.8\%, and improves inference speed by an average of 81.1\% compared to prior diffusion-based methods, while achieving state-of-the-art performance.
format Preprint
id arxiv_https___arxiv_org_abs_2508_21363
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Efficient Diffusion-Based 3D Human Pose Estimation with Hierarchical Temporal Pruning
Bi, Yuquan
Wang, Hongsong
Shi, Xinli
Gui, Zhipeng
Gui, Jie
Tang, Yuan Yan
Computer Vision and Pattern Recognition
Diffusion models have demonstrated strong capabilities in generating high-fidelity 3D human poses, yet their iterative nature and multi-hypothesis requirements incur substantial computational cost. In this paper, we propose an Efficient Diffusion-Based 3D Human Pose Estimation framework with a Hierarchical Temporal Pruning (HTP) strategy, which dynamically prunes redundant pose tokens across both frame and semantic levels while preserving critical motion dynamics. HTP operates in a staged, top-down manner: (1) Temporal Correlation-Enhanced Pruning (TCEP) identifies essential frames by analyzing inter-frame motion correlations through adaptive temporal graph construction; (2) Sparse-Focused Temporal MHSA (SFT MHSA) leverages the resulting frame-level sparsity to reduce attention computation, focusing on motion-relevant tokens; and (3) Mask-Guided Pose Token Pruner (MGPTP) performs fine-grained semantic pruning via clustering, retaining only the most informative pose tokens. Experiments on Human3.6M and MPI-INF-3DHP show that HTP reduces training MACs by 38.5\%, inference MACs by 56.8\%, and improves inference speed by an average of 81.1\% compared to prior diffusion-based methods, while achieving state-of-the-art performance.
title Efficient Diffusion-Based 3D Human Pose Estimation with Hierarchical Temporal Pruning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2508.21363