AnimalMotionCLIP: Embedding motion in CLIP for Animal Behavior Analysis

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhong, Enmin, del-Blanco, Carlos R., Berjón, Daniel, Jaureguizar, Fernando, García, Narciso
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916716055887872
author Zhong, Enmin
del-Blanco, Carlos R.
Berjón, Daniel
Jaureguizar, Fernando
García, Narciso
author_facet Zhong, Enmin
del-Blanco, Carlos R.
Berjón, Daniel
Jaureguizar, Fernando
García, Narciso
contents Recently, there has been a surge of interest in applying deep learning techniques to animal behavior recognition, particularly leveraging pre-trained visual language models, such as CLIP, due to their remarkable generalization capacity across various downstream tasks. However, adapting these models to the specific domain of animal behavior recognition presents two significant challenges: integrating motion information and devising an effective temporal modeling scheme. In this paper, we propose AnimalMotionCLIP to address these challenges by interleaving video frames and optical flow information in the CLIP framework. Additionally, several temporal modeling schemes using an aggregation of classifiers are proposed and compared: dense, semi dense, and sparse. As a result, fine temporal actions can be correctly recognized, which is of vital importance in animal behavior analysis. Experiments on the Animal Kingdom dataset demonstrate that AnimalMotionCLIP achieves superior performance compared to state-of-the-art approaches.
format Preprint
id arxiv_https___arxiv_org_abs_2505_00569
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle AnimalMotionCLIP: Embedding motion in CLIP for Animal Behavior Analysis
Zhong, Enmin
del-Blanco, Carlos R.
Berjón, Daniel
Jaureguizar, Fernando
García, Narciso
Computer Vision and Pattern Recognition
Recently, there has been a surge of interest in applying deep learning techniques to animal behavior recognition, particularly leveraging pre-trained visual language models, such as CLIP, due to their remarkable generalization capacity across various downstream tasks. However, adapting these models to the specific domain of animal behavior recognition presents two significant challenges: integrating motion information and devising an effective temporal modeling scheme. In this paper, we propose AnimalMotionCLIP to address these challenges by interleaving video frames and optical flow information in the CLIP framework. Additionally, several temporal modeling schemes using an aggregation of classifiers are proposed and compared: dense, semi dense, and sparse. As a result, fine temporal actions can be correctly recognized, which is of vital importance in animal behavior analysis. Experiments on the Animal Kingdom dataset demonstrate that AnimalMotionCLIP achieves superior performance compared to state-of-the-art approaches.
title AnimalMotionCLIP: Embedding motion in CLIP for Animal Behavior Analysis
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.00569