Human Pose Estimation from Ambiguous Pressure Recordings with Spatio-temporal Masked Transformers

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Davoodnia, Vandad, Etemad, Ali
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917716838842368
author Davoodnia, Vandad
Etemad, Ali
author_facet Davoodnia, Vandad
Etemad, Ali
contents Despite the impressive performance of vision-based pose estimators, they generally fail to perform well under adverse vision conditions and often don't satisfy the privacy demands of customers. As a result, researchers have begun to study tactile sensing systems as an alternative. However, these systems suffer from noisy and ambiguous recordings. To tackle this problem, we propose a novel solution for pose estimation from ambiguous pressure data. Our method comprises a spatio-temporal vision transformer with an encoder-decoder architecture. Detailed experiments on two popular public datasets reveal that our model outperforms existing solutions in the area. Moreover, we observe that increasing the number of temporal crops in the early stages of the network positively impacts the performance while pre-training the network in a self-supervised setting using a masked auto-encoder approach also further improves the results.
format Preprint
id arxiv_https___arxiv_org_abs_2303_05691
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Human Pose Estimation from Ambiguous Pressure Recordings with Spatio-temporal Masked Transformers
Davoodnia, Vandad
Etemad, Ali
Computer Vision and Pattern Recognition
Machine Learning
Despite the impressive performance of vision-based pose estimators, they generally fail to perform well under adverse vision conditions and often don't satisfy the privacy demands of customers. As a result, researchers have begun to study tactile sensing systems as an alternative. However, these systems suffer from noisy and ambiguous recordings. To tackle this problem, we propose a novel solution for pose estimation from ambiguous pressure data. Our method comprises a spatio-temporal vision transformer with an encoder-decoder architecture. Detailed experiments on two popular public datasets reveal that our model outperforms existing solutions in the area. Moreover, we observe that increasing the number of temporal crops in the early stages of the network positively impacts the performance while pre-training the network in a self-supervised setting using a masked auto-encoder approach also further improves the results.
title Human Pose Estimation from Ambiguous Pressure Recordings with Spatio-temporal Masked Transformers
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2303.05691