Learning the 3D Fauna of the Web

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Zizhang, Litvak, Dor, Li, Ruining, Zhang, Yunzhi, Jakab, Tomas, Rupprecht, Christian, Wu, Shangzhe, Vedaldi, Andrea, Wu, Jiajun
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929297349935104
author Li, Zizhang
Litvak, Dor
Li, Ruining
Zhang, Yunzhi
Jakab, Tomas
Rupprecht, Christian
Wu, Shangzhe
Vedaldi, Andrea
Wu, Jiajun
author_facet Li, Zizhang
Litvak, Dor
Li, Ruining
Zhang, Yunzhi
Jakab, Tomas
Rupprecht, Christian
Wu, Shangzhe
Vedaldi, Andrea
Wu, Jiajun
contents Learning 3D models of all animals on the Earth requires massively scaling up existing solutions. With this ultimate goal in mind, we develop 3D-Fauna, an approach that learns a pan-category deformable 3D animal model for more than 100 animal species jointly. One crucial bottleneck of modeling animals is the limited availability of training data, which we overcome by simply learning from 2D Internet images. We show that prior category-specific attempts fail to generalize to rare species with limited training images. We address this challenge by introducing the Semantic Bank of Skinned Models (SBSM), which automatically discovers a small set of base animal shapes by combining geometric inductive priors with semantic knowledge implicitly captured by an off-the-shelf self-supervised feature extractor. To train such a model, we also contribute a new large-scale dataset of diverse animal species. At inference time, given a single image of any quadruped animal, our model reconstructs an articulated 3D mesh in a feed-forward fashion within seconds.
format Preprint
id arxiv_https___arxiv_org_abs_2401_02400
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Learning the 3D Fauna of the Web
Li, Zizhang
Litvak, Dor
Li, Ruining
Zhang, Yunzhi
Jakab, Tomas
Rupprecht, Christian
Wu, Shangzhe
Vedaldi, Andrea
Wu, Jiajun
Computer Vision and Pattern Recognition
Learning 3D models of all animals on the Earth requires massively scaling up existing solutions. With this ultimate goal in mind, we develop 3D-Fauna, an approach that learns a pan-category deformable 3D animal model for more than 100 animal species jointly. One crucial bottleneck of modeling animals is the limited availability of training data, which we overcome by simply learning from 2D Internet images. We show that prior category-specific attempts fail to generalize to rare species with limited training images. We address this challenge by introducing the Semantic Bank of Skinned Models (SBSM), which automatically discovers a small set of base animal shapes by combining geometric inductive priors with semantic knowledge implicitly captured by an off-the-shelf self-supervised feature extractor. To train such a model, we also contribute a new large-scale dataset of diverse animal species. At inference time, given a single image of any quadruped animal, our model reconstructs an articulated 3D mesh in a feed-forward fashion within seconds.
title Learning the 3D Fauna of the Web
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2401.02400