Contrastive Learning for Multi-Object Tracking with Transformers
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | , , , , |
|---|---|
| Format: | Preprint |
| Publié: |
2023
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866909610645913600 |
|---|---|
| author | De Plaen, Pierre-François Marinello, Nicola Proesmans, Marc Tuytelaars, Tinne Van Gool, Luc |
| author_facet | De Plaen, Pierre-François Marinello, Nicola Proesmans, Marc Tuytelaars, Tinne Van Gool, Luc |
| contents | The DEtection TRansformer (DETR) opened new possibilities for object detection by modeling it as a translation task: converting image features into object-level representations. Previous works typically add expensive modules to DETR to perform Multi-Object Tracking (MOT), resulting in more complicated architectures. We instead show how DETR can be turned into a MOT model by employing an instance-level contrastive loss, a revised sampling strategy and a lightweight assignment method. Our training scheme learns object appearances while preserving detection capabilities and with little overhead. Its performance surpasses the previous state-of-the-art by +2.6 mMOTA on the challenging BDD100K dataset and is comparable to existing transformer-based methods on the MOT17 dataset. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2311_08043 |
| institution | arXiv |
| publishDate | 2023 |
| record_format | arxiv |
| spellingShingle | Contrastive Learning for Multi-Object Tracking with Transformers De Plaen, Pierre-François Marinello, Nicola Proesmans, Marc Tuytelaars, Tinne Van Gool, Luc Computer Vision and Pattern Recognition The DEtection TRansformer (DETR) opened new possibilities for object detection by modeling it as a translation task: converting image features into object-level representations. Previous works typically add expensive modules to DETR to perform Multi-Object Tracking (MOT), resulting in more complicated architectures. We instead show how DETR can be turned into a MOT model by employing an instance-level contrastive loss, a revised sampling strategy and a lightweight assignment method. Our training scheme learns object appearances while preserving detection capabilities and with little overhead. Its performance surpasses the previous state-of-the-art by +2.6 mMOTA on the challenging BDD100K dataset and is comparable to existing transformer-based methods on the MOT17 dataset. |
| title | Contrastive Learning for Multi-Object Tracking with Transformers |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2311.08043 |