Multi-modal Pose Diffuser: A Multimodal Generative Conditional Pose Prior

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ta, Calvin-Khang, Dutta, Arindam, Kundu, Rohit, Lal, Rohit, Cruz, Hannah Dela, Raychaudhuri, Dripta S., Roy-Chowdhury, Amit
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909354217701376
author Ta, Calvin-Khang
Dutta, Arindam
Kundu, Rohit
Lal, Rohit
Cruz, Hannah Dela
Raychaudhuri, Dripta S.
Roy-Chowdhury, Amit
author_facet Ta, Calvin-Khang
Dutta, Arindam
Kundu, Rohit
Lal, Rohit
Cruz, Hannah Dela
Raychaudhuri, Dripta S.
Roy-Chowdhury, Amit
contents The Skinned Multi-Person Linear (SMPL) model plays a crucial role in 3D human pose estimation, providing a streamlined yet effective representation of the human body. However, ensuring the validity of SMPL configurations during tasks such as human mesh regression remains a significant challenge , highlighting the necessity for a robust human pose prior capable of discerning realistic human poses. To address this, we introduce MOPED: \underline{M}ulti-m\underline{O}dal \underline{P}os\underline{E} \underline{D}iffuser. MOPED is the first method to leverage a novel multi-modal conditional diffusion model as a prior for SMPL pose parameters. Our method offers powerful unconditional pose generation with the ability to condition on multi-modal inputs such as images and text. This capability enhances the applicability of our approach by incorporating additional context often overlooked in traditional pose priors. Extensive experiments across three distinct tasks-pose estimation, pose denoising, and pose completion-demonstrate that our multi-modal diffusion model-based prior significantly outperforms existing methods. These results indicate that our model captures a broader spectrum of plausible human poses.
format Preprint
id arxiv_https___arxiv_org_abs_2410_14540
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Multi-modal Pose Diffuser: A Multimodal Generative Conditional Pose Prior
Ta, Calvin-Khang
Dutta, Arindam
Kundu, Rohit
Lal, Rohit
Cruz, Hannah Dela
Raychaudhuri, Dripta S.
Roy-Chowdhury, Amit
Computer Vision and Pattern Recognition
The Skinned Multi-Person Linear (SMPL) model plays a crucial role in 3D human pose estimation, providing a streamlined yet effective representation of the human body. However, ensuring the validity of SMPL configurations during tasks such as human mesh regression remains a significant challenge , highlighting the necessity for a robust human pose prior capable of discerning realistic human poses. To address this, we introduce MOPED: \underline{M}ulti-m\underline{O}dal \underline{P}os\underline{E} \underline{D}iffuser. MOPED is the first method to leverage a novel multi-modal conditional diffusion model as a prior for SMPL pose parameters. Our method offers powerful unconditional pose generation with the ability to condition on multi-modal inputs such as images and text. This capability enhances the applicability of our approach by incorporating additional context often overlooked in traditional pose priors. Extensive experiments across three distinct tasks-pose estimation, pose denoising, and pose completion-demonstrate that our multi-modal diffusion model-based prior significantly outperforms existing methods. These results indicate that our model captures a broader spectrum of plausible human poses.
title Multi-modal Pose Diffuser: A Multimodal Generative Conditional Pose Prior
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2410.14540