SSD-Poser: Avatar Pose Estimation with State Space Duality from Sparse Observations

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhao, Shuting, Bai, Linxin, Shao, Liangjing, Zhang, Ye, Chen, Xinrong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910918983548928
author Zhao, Shuting
Bai, Linxin
Shao, Liangjing
Zhang, Ye
Chen, Xinrong
author_facet Zhao, Shuting
Bai, Linxin
Shao, Liangjing
Zhang, Ye
Chen, Xinrong
contents The growing applications of AR/VR increase the demand for real-time full-body pose estimation from Head-Mounted Displays (HMDs). Although HMDs provide joint signals from the head and hands, reconstructing a full-body pose remains challenging due to the unconstrained lower body. Recent advancements often rely on conventional neural networks and generative models to improve performance in this task, such as Transformers and diffusion models. However, these approaches struggle to strike a balance between achieving precise pose reconstruction and maintaining fast inference speed. To overcome these challenges, a lightweight and efficient model, SSD-Poser, is designed for robust full-body motion estimation from sparse observations. SSD-Poser incorporates a well-designed hybrid encoder, State Space Attention Encoders, to adapt the state space duality to complex motion poses and enable real-time realistic pose reconstruction. Moreover, a Frequency-Aware Decoder is introduced to mitigate jitter caused by variable-frequency motion signals, remarkably enhancing the motion smoothness. Comprehensive experiments on the AMASS dataset demonstrate that SSD-Poser achieves exceptional accuracy and computational efficiency, showing outstanding inference efficiency compared to state-of-the-art methods.
format Preprint
id arxiv_https___arxiv_org_abs_2504_18332
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SSD-Poser: Avatar Pose Estimation with State Space Duality from Sparse Observations
Zhao, Shuting
Bai, Linxin
Shao, Liangjing
Zhang, Ye
Chen, Xinrong
Computer Vision and Pattern Recognition
Human-Computer Interaction
68U05
The growing applications of AR/VR increase the demand for real-time full-body pose estimation from Head-Mounted Displays (HMDs). Although HMDs provide joint signals from the head and hands, reconstructing a full-body pose remains challenging due to the unconstrained lower body. Recent advancements often rely on conventional neural networks and generative models to improve performance in this task, such as Transformers and diffusion models. However, these approaches struggle to strike a balance between achieving precise pose reconstruction and maintaining fast inference speed. To overcome these challenges, a lightweight and efficient model, SSD-Poser, is designed for robust full-body motion estimation from sparse observations. SSD-Poser incorporates a well-designed hybrid encoder, State Space Attention Encoders, to adapt the state space duality to complex motion poses and enable real-time realistic pose reconstruction. Moreover, a Frequency-Aware Decoder is introduced to mitigate jitter caused by variable-frequency motion signals, remarkably enhancing the motion smoothness. Comprehensive experiments on the AMASS dataset demonstrate that SSD-Poser achieves exceptional accuracy and computational efficiency, showing outstanding inference efficiency compared to state-of-the-art methods.
title SSD-Poser: Avatar Pose Estimation with State Space Duality from Sparse Observations
topic Computer Vision and Pattern Recognition
Human-Computer Interaction
68U05
url https://arxiv.org/abs/2504.18332