Deep Non-rigid Structure-from-Motion: A Sequence-to-Sequence Translation Perspective

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Deng, Hui, Zhang, Tong, Dai, Yuchao, Shi, Jiawei, Zhong, Yiran, Li, Hongdong
Format: Preprint
Published: 2022
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914910375510016
author Deng, Hui
Zhang, Tong
Dai, Yuchao
Shi, Jiawei
Zhong, Yiran
Li, Hongdong
author_facet Deng, Hui
Zhang, Tong
Dai, Yuchao
Shi, Jiawei
Zhong, Yiran
Li, Hongdong
contents Directly regressing the non-rigid shape and camera pose from the individual 2D frame is ill-suited to the Non-Rigid Structure-from-Motion (NRSfM) problem. This frame-by-frame 3D reconstruction pipeline overlooks the inherent spatial-temporal nature of NRSfM, i.e., reconstructing the whole 3D sequence from the input 2D sequence. In this paper, we propose to model deep NRSfM from a sequence-to-sequence translation perspective, where the input 2D frame sequence is taken as a whole to reconstruct the deforming 3D non-rigid shape sequence. First, we apply a shape-motion predictor to estimate the initial non-rigid shape and camera motion from a single frame. Then we propose a context modeling module to model camera motions and complex non-rigid shapes. To tackle the difficulty in enforcing the global structure constraint within the deep framework, we propose to impose the union-of-subspace structure by replacing the self-expressiveness layer with multi-head attention and delayed regularizers, which enables end-to-end batch-wise training. Experimental results across different datasets such as Human3.6M, CMU Mocap and InterHand prove the superiority of our framework.
format Preprint
id arxiv_https___arxiv_org_abs_2204_04730
institution arXiv
publishDate 2022
record_format arxiv
spellingShingle Deep Non-rigid Structure-from-Motion: A Sequence-to-Sequence Translation Perspective
Deng, Hui
Zhang, Tong
Dai, Yuchao
Shi, Jiawei
Zhong, Yiran
Li, Hongdong
Computer Vision and Pattern Recognition
Directly regressing the non-rigid shape and camera pose from the individual 2D frame is ill-suited to the Non-Rigid Structure-from-Motion (NRSfM) problem. This frame-by-frame 3D reconstruction pipeline overlooks the inherent spatial-temporal nature of NRSfM, i.e., reconstructing the whole 3D sequence from the input 2D sequence. In this paper, we propose to model deep NRSfM from a sequence-to-sequence translation perspective, where the input 2D frame sequence is taken as a whole to reconstruct the deforming 3D non-rigid shape sequence. First, we apply a shape-motion predictor to estimate the initial non-rigid shape and camera motion from a single frame. Then we propose a context modeling module to model camera motions and complex non-rigid shapes. To tackle the difficulty in enforcing the global structure constraint within the deep framework, we propose to impose the union-of-subspace structure by replacing the self-expressiveness layer with multi-head attention and delayed regularizers, which enables end-to-end batch-wise training. Experimental results across different datasets such as Human3.6M, CMU Mocap and InterHand prove the superiority of our framework.
title Deep Non-rigid Structure-from-Motion: A Sequence-to-Sequence Translation Perspective
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2204.04730