RUMPL: Ray-Based Transformers for Universal Multi-View 2D to 3D Human Pose Lifting

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ghasemzadeh, Seyed Abolfazl, Alahi, Alexandre, De Vleeschouwer, Christophe
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908718288863232
author Ghasemzadeh, Seyed Abolfazl
Alahi, Alexandre
De Vleeschouwer, Christophe
author_facet Ghasemzadeh, Seyed Abolfazl
Alahi, Alexandre
De Vleeschouwer, Christophe
contents Estimating 3D human poses from 2D images remains challenging due to occlusions and projective ambiguity. Multi-view learning-based approaches mitigate these issues but often fail to generalize to real-world scenarios, as large-scale multi-view datasets with 3D ground truth are scarce and captured under constrained conditions. To overcome this limitation, recent methods rely on 2D pose estimation combined with 2D-to-3D pose lifting trained on synthetic data. Building on our previous MPL framework, we propose RUMPL, a transformer-based 3D pose lifter that introduces a 3D ray-based representation of 2D keypoints. This formulation makes the model independent of camera calibration and the number of views, enabling universal deployment across arbitrary multi-view configurations without retraining or fine-tuning. A new View Fusion Transformer leverages learned fused-ray tokens to aggregate information along rays, further improving multi-view consistency. Extensive experiments demonstrate that RUMPL reduces MPJPE by up to 53% compared to triangulation and over 60% compared to transformer-based image-representation baselines. Results on new benchmarks, including in-the-wild multi-view and multi-person datasets, confirm its robustness and scalability. The framework's source code is available at https://github.com/aghasemzadeh/OpenRUMPL
format Preprint
id arxiv_https___arxiv_org_abs_2512_15488
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle RUMPL: Ray-Based Transformers for Universal Multi-View 2D to 3D Human Pose Lifting
Ghasemzadeh, Seyed Abolfazl
Alahi, Alexandre
De Vleeschouwer, Christophe
Computer Vision and Pattern Recognition
Estimating 3D human poses from 2D images remains challenging due to occlusions and projective ambiguity. Multi-view learning-based approaches mitigate these issues but often fail to generalize to real-world scenarios, as large-scale multi-view datasets with 3D ground truth are scarce and captured under constrained conditions. To overcome this limitation, recent methods rely on 2D pose estimation combined with 2D-to-3D pose lifting trained on synthetic data. Building on our previous MPL framework, we propose RUMPL, a transformer-based 3D pose lifter that introduces a 3D ray-based representation of 2D keypoints. This formulation makes the model independent of camera calibration and the number of views, enabling universal deployment across arbitrary multi-view configurations without retraining or fine-tuning. A new View Fusion Transformer leverages learned fused-ray tokens to aggregate information along rays, further improving multi-view consistency. Extensive experiments demonstrate that RUMPL reduces MPJPE by up to 53% compared to triangulation and over 60% compared to transformer-based image-representation baselines. Results on new benchmarks, including in-the-wild multi-view and multi-person datasets, confirm its robustness and scalability. The framework's source code is available at https://github.com/aghasemzadeh/OpenRUMPL
title RUMPL: Ray-Based Transformers for Universal Multi-View 2D to 3D Human Pose Lifting
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.15488