Exploiting Spatiotemporal Properties for Efficient Event-Driven Human Pose Estimation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhou, Haoxian, Xu, Chuanzhi, Chen, Langyi, Ye, Pengfei, Chen, Haodong, Chung, Yuk Ying, Qu, Qiang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911491676962816
author Zhou, Haoxian
Xu, Chuanzhi
Chen, Langyi
Ye, Pengfei
Chen, Haodong
Chung, Yuk Ying
Qu, Qiang
author_facet Zhou, Haoxian
Xu, Chuanzhi
Chen, Langyi
Ye, Pengfei
Chen, Haodong
Chung, Yuk Ying
Qu, Qiang
contents Human pose estimation focuses on predicting body keypoints to analyze human motion. Currently, most pose estimation tasks rely on conventional RGB cameras. In contrast, event cameras provide high temporal resolution and low latency, enabling robust estimation under challenging conditions and opening up new possibilities for pose estimation. However, most existing methods convert event streams into dense event frames, which adds extra computation and sacrifices the high temporal resolution of the event signal. In this work, we aim to exploit the spatiotemporal properties of event streams based on point cloud-based framework, designed to enhance human pose estimation performance while maintaining computational efficiency. We design Event Temporal Slicing Convolution module to capture short-term dependencies across event slices, and combine it with Event Slice Sequencing module for structured temporal modeling. We further propose an edge-enhanced point cloud-based event representation to enhance spatial edge information under sparse event conditions to further improve performance. Experiments on the DHP19 dataset show our proposed method consistently improves performance across three representative point cloud backbones: PointNet, DGCNN, and Point Transformer, with an average MPJPE reduction of 4%.
format Preprint
id arxiv_https___arxiv_org_abs_2512_06306
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Exploiting Spatiotemporal Properties for Efficient Event-Driven Human Pose Estimation
Zhou, Haoxian
Xu, Chuanzhi
Chen, Langyi
Ye, Pengfei
Chen, Haodong
Chung, Yuk Ying
Qu, Qiang
Computer Vision and Pattern Recognition
Artificial Intelligence
Human pose estimation focuses on predicting body keypoints to analyze human motion. Currently, most pose estimation tasks rely on conventional RGB cameras. In contrast, event cameras provide high temporal resolution and low latency, enabling robust estimation under challenging conditions and opening up new possibilities for pose estimation. However, most existing methods convert event streams into dense event frames, which adds extra computation and sacrifices the high temporal resolution of the event signal. In this work, we aim to exploit the spatiotemporal properties of event streams based on point cloud-based framework, designed to enhance human pose estimation performance while maintaining computational efficiency. We design Event Temporal Slicing Convolution module to capture short-term dependencies across event slices, and combine it with Event Slice Sequencing module for structured temporal modeling. We further propose an edge-enhanced point cloud-based event representation to enhance spatial edge information under sparse event conditions to further improve performance. Experiments on the DHP19 dataset show our proposed method consistently improves performance across three representative point cloud backbones: PointNet, DGCNN, and Point Transformer, with an average MPJPE reduction of 4%.
title Exploiting Spatiotemporal Properties for Efficient Event-Driven Human Pose Estimation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2512.06306