Contrastive Learning for Multi-Object Tracking with Transformers

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: De Plaen, Pierre-François, Marinello, Nicola, Proesmans, Marc, Tuytelaars, Tinne, Van Gool, Luc
Format: Preprint
Publié: 2023
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866909610645913600
author De Plaen, Pierre-François
Marinello, Nicola
Proesmans, Marc
Tuytelaars, Tinne
Van Gool, Luc
author_facet De Plaen, Pierre-François
Marinello, Nicola
Proesmans, Marc
Tuytelaars, Tinne
Van Gool, Luc
contents The DEtection TRansformer (DETR) opened new possibilities for object detection by modeling it as a translation task: converting image features into object-level representations. Previous works typically add expensive modules to DETR to perform Multi-Object Tracking (MOT), resulting in more complicated architectures. We instead show how DETR can be turned into a MOT model by employing an instance-level contrastive loss, a revised sampling strategy and a lightweight assignment method. Our training scheme learns object appearances while preserving detection capabilities and with little overhead. Its performance surpasses the previous state-of-the-art by +2.6 mMOTA on the challenging BDD100K dataset and is comparable to existing transformer-based methods on the MOT17 dataset.
format Preprint
id arxiv_https___arxiv_org_abs_2311_08043
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Contrastive Learning for Multi-Object Tracking with Transformers
De Plaen, Pierre-François
Marinello, Nicola
Proesmans, Marc
Tuytelaars, Tinne
Van Gool, Luc
Computer Vision and Pattern Recognition
The DEtection TRansformer (DETR) opened new possibilities for object detection by modeling it as a translation task: converting image features into object-level representations. Previous works typically add expensive modules to DETR to perform Multi-Object Tracking (MOT), resulting in more complicated architectures. We instead show how DETR can be turned into a MOT model by employing an instance-level contrastive loss, a revised sampling strategy and a lightweight assignment method. Our training scheme learns object appearances while preserving detection capabilities and with little overhead. Its performance surpasses the previous state-of-the-art by +2.6 mMOTA on the challenging BDD100K dataset and is comparable to existing transformer-based methods on the MOT17 dataset.
title Contrastive Learning for Multi-Object Tracking with Transformers
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2311.08043