CoMotion: Concurrent Multi-person 3D Motion

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Newell, Alejandro, Hu, Peiyun, Lipson, Lahav, Richter, Stephan R., Koltun, Vladlen
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917986981380096
author Newell, Alejandro
Hu, Peiyun
Lipson, Lahav
Richter, Stephan R.
Koltun, Vladlen
author_facet Newell, Alejandro
Hu, Peiyun
Lipson, Lahav
Richter, Stephan R.
Koltun, Vladlen
contents We introduce an approach for detecting and tracking detailed 3D poses of multiple people from a single monocular camera stream. Our system maintains temporally coherent predictions in crowded scenes filled with difficult poses and occlusions. Our model performs both strong per-frame detection and a learned pose update to track people from frame to frame. Rather than match detections across time, poses are updated directly from a new input image, which enables online tracking through occlusion. We train on numerous image and video datasets leveraging pseudo-labeled annotations to produce a model that matches state-of-the-art systems in 3D pose estimation accuracy while being faster and more accurate in tracking multiple people through time. Code and weights are provided at https://github.com/apple/ml-comotion
format Preprint
id arxiv_https___arxiv_org_abs_2504_12186
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CoMotion: Concurrent Multi-person 3D Motion
Newell, Alejandro
Hu, Peiyun
Lipson, Lahav
Richter, Stephan R.
Koltun, Vladlen
Computer Vision and Pattern Recognition
Machine Learning
We introduce an approach for detecting and tracking detailed 3D poses of multiple people from a single monocular camera stream. Our system maintains temporally coherent predictions in crowded scenes filled with difficult poses and occlusions. Our model performs both strong per-frame detection and a learned pose update to track people from frame to frame. Rather than match detections across time, poses are updated directly from a new input image, which enables online tracking through occlusion. We train on numerous image and video datasets leveraging pseudo-labeled annotations to produce a model that matches state-of-the-art systems in 3D pose estimation accuracy while being faster and more accurate in tracking multiple people through time. Code and weights are provided at https://github.com/apple/ml-comotion
title CoMotion: Concurrent Multi-person 3D Motion
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2504.12186