Object Agnostic 3D Lifting in Space and Time

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fusco, Christopher, Ch'ng, Shin-Fang, Dabhi, Mosam, Lucey, Simon
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910819769384960
author Fusco, Christopher
Ch'ng, Shin-Fang
Dabhi, Mosam
Lucey, Simon
author_facet Fusco, Christopher
Ch'ng, Shin-Fang
Dabhi, Mosam
Lucey, Simon
contents We present a spatio-temporal perspective on category-agnostic 3D lifting of 2D keypoints over a temporal sequence. Our approach differs from existing state-of-the-art methods that are either: (i) object-agnostic, but can only operate on individual frames, or (ii) can model space-time dependencies, but are only designed to work with a single object category. Our approach is grounded in two core principles. First, general information about similar objects can be leveraged to achieve better performance when there is little object-specific training data. Second, a temporally-proximate context window is advantageous for achieving consistency throughout a sequence. These two principles allow us to outperform current state-of-the-art methods on per-frame and per-sequence metrics for a variety of animal categories. Lastly, we release a new synthetic dataset containing 3D skeletons and motion sequences for a variety of animal categories.
format Preprint
id arxiv_https___arxiv_org_abs_2412_01166
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Object Agnostic 3D Lifting in Space and Time
Fusco, Christopher
Ch'ng, Shin-Fang
Dabhi, Mosam
Lucey, Simon
Computer Vision and Pattern Recognition
Artificial Intelligence
We present a spatio-temporal perspective on category-agnostic 3D lifting of 2D keypoints over a temporal sequence. Our approach differs from existing state-of-the-art methods that are either: (i) object-agnostic, but can only operate on individual frames, or (ii) can model space-time dependencies, but are only designed to work with a single object category. Our approach is grounded in two core principles. First, general information about similar objects can be leveraged to achieve better performance when there is little object-specific training data. Second, a temporally-proximate context window is advantageous for achieving consistency throughout a sequence. These two principles allow us to outperform current state-of-the-art methods on per-frame and per-sequence metrics for a variety of animal categories. Lastly, we release a new synthetic dataset containing 3D skeletons and motion sequences for a variety of animal categories.
title Object Agnostic 3D Lifting in Space and Time
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2412.01166