Domain-Adaptive Pretraining Improves Primate Behavior Recognition

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mueller, Felix B., Lueddecke, Timo, Vogg, Richard, Ecker, Alexander S.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908540133703680
author Mueller, Felix B.
Lueddecke, Timo
Vogg, Richard
Ecker, Alexander S.
author_facet Mueller, Felix B.
Lueddecke, Timo
Vogg, Richard
Ecker, Alexander S.
contents Computer vision for animal behavior offers promising tools to aid research in ecology, cognition, and to support conservation efforts. Video camera traps allow for large-scale data collection, but high labeling costs remain a bottleneck to creating large-scale datasets. We thus need data-efficient learning approaches. In this work, we show that we can utilize self-supervised learning to considerably improve action recognition on primate behavior. On two datasets of great ape behavior (PanAf and ChimpACT), we outperform published state-of-the-art action recognition models by 6.1 %pt. accuracy and 6.3 %pt. mAP, respectively. We achieve this by utilizing a pretrained V-JEPA model and applying domain-adaptive pretraining (DAP), i.e. continuing the pretraining with in-domain data. We show that most of the performance gain stems from the DAP. Our method promises great potential for improving the recognition of animal behavior, as DAP does not require labeled samples. Code is available at https://github.com/ecker-lab/dap-behavior
format Preprint
id arxiv_https___arxiv_org_abs_2509_12193
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Domain-Adaptive Pretraining Improves Primate Behavior Recognition
Mueller, Felix B.
Lueddecke, Timo
Vogg, Richard
Ecker, Alexander S.
Computer Vision and Pattern Recognition
I.4.8; I.2.10; I.5
Computer vision for animal behavior offers promising tools to aid research in ecology, cognition, and to support conservation efforts. Video camera traps allow for large-scale data collection, but high labeling costs remain a bottleneck to creating large-scale datasets. We thus need data-efficient learning approaches. In this work, we show that we can utilize self-supervised learning to considerably improve action recognition on primate behavior. On two datasets of great ape behavior (PanAf and ChimpACT), we outperform published state-of-the-art action recognition models by 6.1 %pt. accuracy and 6.3 %pt. mAP, respectively. We achieve this by utilizing a pretrained V-JEPA model and applying domain-adaptive pretraining (DAP), i.e. continuing the pretraining with in-domain data. We show that most of the performance gain stems from the DAP. Our method promises great potential for improving the recognition of animal behavior, as DAP does not require labeled samples. Code is available at https://github.com/ecker-lab/dap-behavior
title Domain-Adaptive Pretraining Improves Primate Behavior Recognition
topic Computer Vision and Pattern Recognition
I.4.8; I.2.10; I.5
url https://arxiv.org/abs/2509.12193