DETRPose: Real-time end-to-end transformer model for multi-person pose estimation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Janampa, Sebastian, Pattichis, Marios
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915346230804480
author Janampa, Sebastian
Pattichis, Marios
author_facet Janampa, Sebastian
Pattichis, Marios
contents Multi-person pose estimation (MPPE) estimates keypoints for all individuals present in an image. MPPE is a fundamental task for several applications in computer vision and virtual reality. Unfortunately, there are currently no transformer-based models that can perform MPPE in real time. The paper presents a family of transformer-based models capable of performing multi-person 2D pose estimation in real-time. Our approach utilizes a modified decoder architecture and keypoint similarity metrics to generate both positive and negative queries, thereby enhancing the quality of the selected queries within the architecture. Compared to state-of-the-art models, our proposed models train much faster, using 5 to 10 times fewer epochs, with competitive inference times without requiring quantization libraries to speed up the model. Furthermore, our proposed models provide competitive results or outperform alternative models, often using significantly fewer parameters.
format Preprint
id arxiv_https___arxiv_org_abs_2506_13027
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DETRPose: Real-time end-to-end transformer model for multi-person pose estimation
Janampa, Sebastian
Pattichis, Marios
Computer Vision and Pattern Recognition
Multi-person pose estimation (MPPE) estimates keypoints for all individuals present in an image. MPPE is a fundamental task for several applications in computer vision and virtual reality. Unfortunately, there are currently no transformer-based models that can perform MPPE in real time. The paper presents a family of transformer-based models capable of performing multi-person 2D pose estimation in real-time. Our approach utilizes a modified decoder architecture and keypoint similarity metrics to generate both positive and negative queries, thereby enhancing the quality of the selected queries within the architecture. Compared to state-of-the-art models, our proposed models train much faster, using 5 to 10 times fewer epochs, with competitive inference times without requiring quantization libraries to speed up the model. Furthermore, our proposed models provide competitive results or outperform alternative models, often using significantly fewer parameters.
title DETRPose: Real-time end-to-end transformer model for multi-person pose estimation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.13027