Learning the 3D Fauna of the Web
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866929297349935104 |
|---|---|
| author | Li, Zizhang Litvak, Dor Li, Ruining Zhang, Yunzhi Jakab, Tomas Rupprecht, Christian Wu, Shangzhe Vedaldi, Andrea Wu, Jiajun |
| author_facet | Li, Zizhang Litvak, Dor Li, Ruining Zhang, Yunzhi Jakab, Tomas Rupprecht, Christian Wu, Shangzhe Vedaldi, Andrea Wu, Jiajun |
| contents | Learning 3D models of all animals on the Earth requires massively scaling up existing solutions. With this ultimate goal in mind, we develop 3D-Fauna, an approach that learns a pan-category deformable 3D animal model for more than 100 animal species jointly. One crucial bottleneck of modeling animals is the limited availability of training data, which we overcome by simply learning from 2D Internet images. We show that prior category-specific attempts fail to generalize to rare species with limited training images. We address this challenge by introducing the Semantic Bank of Skinned Models (SBSM), which automatically discovers a small set of base animal shapes by combining geometric inductive priors with semantic knowledge implicitly captured by an off-the-shelf self-supervised feature extractor. To train such a model, we also contribute a new large-scale dataset of diverse animal species. At inference time, given a single image of any quadruped animal, our model reconstructs an articulated 3D mesh in a feed-forward fashion within seconds. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2401_02400 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Learning the 3D Fauna of the Web Li, Zizhang Litvak, Dor Li, Ruining Zhang, Yunzhi Jakab, Tomas Rupprecht, Christian Wu, Shangzhe Vedaldi, Andrea Wu, Jiajun Computer Vision and Pattern Recognition Learning 3D models of all animals on the Earth requires massively scaling up existing solutions. With this ultimate goal in mind, we develop 3D-Fauna, an approach that learns a pan-category deformable 3D animal model for more than 100 animal species jointly. One crucial bottleneck of modeling animals is the limited availability of training data, which we overcome by simply learning from 2D Internet images. We show that prior category-specific attempts fail to generalize to rare species with limited training images. We address this challenge by introducing the Semantic Bank of Skinned Models (SBSM), which automatically discovers a small set of base animal shapes by combining geometric inductive priors with semantic knowledge implicitly captured by an off-the-shelf self-supervised feature extractor. To train such a model, we also contribute a new large-scale dataset of diverse animal species. At inference time, given a single image of any quadruped animal, our model reconstructs an articulated 3D mesh in a feed-forward fashion within seconds. |
| title | Learning the 3D Fauna of the Web |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2401.02400 |