Beyond Boxes: Mask-Guided Spatio-Temporal Feature Aggregation for Video Object Detection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hashmi, Khurram Azeem, Sheikh, Talha Uddin, Stricker, Didier, Afzal, Muhammad Zeshan
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916510760435712
author Hashmi, Khurram Azeem
Sheikh, Talha Uddin
Stricker, Didier
Afzal, Muhammad Zeshan
author_facet Hashmi, Khurram Azeem
Sheikh, Talha Uddin
Stricker, Didier
Afzal, Muhammad Zeshan
contents The primary challenge in Video Object Detection (VOD) is effectively exploiting temporal information to enhance object representations. Traditional strategies, such as aggregating region proposals, often suffer from feature variance due to the inclusion of background information. We introduce a novel instance mask-based feature aggregation approach, significantly refining this process and deepening the understanding of object dynamics across video frames. We present FAIM, a new VOD method that enhances temporal Feature Aggregation by leveraging Instance Mask features. In particular, we propose the lightweight Instance Feature Extraction Module (IFEM) to learn instance mask features and the Temporal Instance Classification Aggregation Module (TICAM) to aggregate instance mask and classification features across video frames. Using YOLOX as a base detector, FAIM achieves 87.9% mAP on the ImageNet VID dataset at 33 FPS on a single 2080Ti GPU, setting a new benchmark for the speed-accuracy trade-off. Additional experiments on multiple datasets validate that our approach is robust, method-agnostic, and effective in multi-object tracking, demonstrating its broader applicability to video understanding tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2412_04915
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Beyond Boxes: Mask-Guided Spatio-Temporal Feature Aggregation for Video Object Detection
Hashmi, Khurram Azeem
Sheikh, Talha Uddin
Stricker, Didier
Afzal, Muhammad Zeshan
Computer Vision and Pattern Recognition
The primary challenge in Video Object Detection (VOD) is effectively exploiting temporal information to enhance object representations. Traditional strategies, such as aggregating region proposals, often suffer from feature variance due to the inclusion of background information. We introduce a novel instance mask-based feature aggregation approach, significantly refining this process and deepening the understanding of object dynamics across video frames. We present FAIM, a new VOD method that enhances temporal Feature Aggregation by leveraging Instance Mask features. In particular, we propose the lightweight Instance Feature Extraction Module (IFEM) to learn instance mask features and the Temporal Instance Classification Aggregation Module (TICAM) to aggregate instance mask and classification features across video frames. Using YOLOX as a base detector, FAIM achieves 87.9% mAP on the ImageNet VID dataset at 33 FPS on a single 2080Ti GPU, setting a new benchmark for the speed-accuracy trade-off. Additional experiments on multiple datasets validate that our approach is robust, method-agnostic, and effective in multi-object tracking, demonstrating its broader applicability to video understanding tasks.
title Beyond Boxes: Mask-Guided Spatio-Temporal Feature Aggregation for Video Object Detection
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.04915