AniMer: Animal Pose and Shape Estimation Using Family Aware Transformer

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Lyu, Jin, Zhu, Tianyi, Gu, Yi, Lin, Li, Cheng, Pujin, Liu, Yebin, Tang, Xiaoying, An, Liang
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915370990829568
author Lyu, Jin
Zhu, Tianyi
Gu, Yi
Lin, Li
Cheng, Pujin
Liu, Yebin
Tang, Xiaoying
An, Liang
author_facet Lyu, Jin
Zhu, Tianyi
Gu, Yi
Lin, Li
Cheng, Pujin
Liu, Yebin
Tang, Xiaoying
An, Liang
contents Quantitative analysis of animal behavior and biomechanics requires accurate animal pose and shape estimation across species, and is important for animal welfare and biological research. However, the small network capacity of previous methods and limited multi-species dataset leave this problem underexplored. To this end, this paper presents AniMer to estimate animal pose and shape using family aware Transformer, enhancing the reconstruction accuracy of diverse quadrupedal families. A key insight of AniMer is its integration of a high-capacity Transformer-based backbone and an animal family supervised contrastive learning scheme, unifying the discriminative understanding of various quadrupedal shapes within a single framework. For effective training, we aggregate most available open-sourced quadrupedal datasets, either with 3D or 2D labels. To improve the diversity of 3D labeled data, we introduce CtrlAni3D, a novel large-scale synthetic dataset created through a new diffusion-based conditional image generation pipeline. CtrlAni3D consists of about 10k images with pixel-aligned SMAL labels. In total, we obtain 41.3k annotated images for training and validation. Consequently, the combination of a family aware Transformer network and an expansive dataset enables AniMer to outperform existing methods not only on 3D datasets like Animal3D and CtrlAni3D, but also on out-of-distribution Animal Kingdom dataset. Ablation studies further demonstrate the effectiveness of our network design and CtrlAni3D in enhancing the performance of AniMer for in-the-wild applications. The project page of AniMer is https://luoxue-star.github.io/AniMer_project_page/.
format Preprint
id arxiv_https___arxiv_org_abs_2412_00837
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle AniMer: Animal Pose and Shape Estimation Using Family Aware Transformer
Lyu, Jin
Zhu, Tianyi
Gu, Yi
Lin, Li
Cheng, Pujin
Liu, Yebin
Tang, Xiaoying
An, Liang
Computer Vision and Pattern Recognition
Quantitative analysis of animal behavior and biomechanics requires accurate animal pose and shape estimation across species, and is important for animal welfare and biological research. However, the small network capacity of previous methods and limited multi-species dataset leave this problem underexplored. To this end, this paper presents AniMer to estimate animal pose and shape using family aware Transformer, enhancing the reconstruction accuracy of diverse quadrupedal families. A key insight of AniMer is its integration of a high-capacity Transformer-based backbone and an animal family supervised contrastive learning scheme, unifying the discriminative understanding of various quadrupedal shapes within a single framework. For effective training, we aggregate most available open-sourced quadrupedal datasets, either with 3D or 2D labels. To improve the diversity of 3D labeled data, we introduce CtrlAni3D, a novel large-scale synthetic dataset created through a new diffusion-based conditional image generation pipeline. CtrlAni3D consists of about 10k images with pixel-aligned SMAL labels. In total, we obtain 41.3k annotated images for training and validation. Consequently, the combination of a family aware Transformer network and an expansive dataset enables AniMer to outperform existing methods not only on 3D datasets like Animal3D and CtrlAni3D, but also on out-of-distribution Animal Kingdom dataset. Ablation studies further demonstrate the effectiveness of our network design and CtrlAni3D in enhancing the performance of AniMer for in-the-wild applications. The project page of AniMer is https://luoxue-star.github.io/AniMer_project_page/.
title AniMer: Animal Pose and Shape Estimation Using Family Aware Transformer
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.00837