FastHMR: Accelerating Human Mesh Recovery via Token and Layer Merging with Diffusion Decoding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mehraban, Soroush, Iaboni, Andrea, Taati, Babak
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917007123808256
author Mehraban, Soroush
Iaboni, Andrea
Taati, Babak
author_facet Mehraban, Soroush
Iaboni, Andrea
Taati, Babak
contents Recent transformer-based models for 3D Human Mesh Recovery (HMR) have achieved strong performance but often suffer from high computational cost and complexity due to deep transformer architectures and redundant tokens. In this paper, we introduce two HMR-specific merging strategies: Error-Constrained Layer Merging (ECLM) and Mask-guided Token Merging (Mask-ToMe). ECLM selectively merges transformer layers that have minimal impact on the Mean Per Joint Position Error (MPJPE), while Mask-ToMe focuses on merging background tokens that contribute little to the final prediction. To further address the potential performance drop caused by merging, we propose a diffusion-based decoder that incorporates temporal context and leverages pose priors learned from large-scale motion capture datasets. Experiments across multiple benchmarks demonstrate that our method achieves up to 2.3x speed-up while slightly improving performance over the baseline.
format Preprint
id arxiv_https___arxiv_org_abs_2510_10868
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle FastHMR: Accelerating Human Mesh Recovery via Token and Layer Merging with Diffusion Decoding
Mehraban, Soroush
Iaboni, Andrea
Taati, Babak
Computer Vision and Pattern Recognition
Recent transformer-based models for 3D Human Mesh Recovery (HMR) have achieved strong performance but often suffer from high computational cost and complexity due to deep transformer architectures and redundant tokens. In this paper, we introduce two HMR-specific merging strategies: Error-Constrained Layer Merging (ECLM) and Mask-guided Token Merging (Mask-ToMe). ECLM selectively merges transformer layers that have minimal impact on the Mean Per Joint Position Error (MPJPE), while Mask-ToMe focuses on merging background tokens that contribute little to the final prediction. To further address the potential performance drop caused by merging, we propose a diffusion-based decoder that incorporates temporal context and leverages pose priors learned from large-scale motion capture datasets. Experiments across multiple benchmarks demonstrate that our method achieves up to 2.3x speed-up while slightly improving performance over the baseline.
title FastHMR: Accelerating Human Mesh Recovery via Token and Layer Merging with Diffusion Decoding
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.10868