DecoderTracker: Decoder-Only Method for Multiple-Object Tracking

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Pan, Liao, Feng, Yang, Wenhui, Zhao, Jinwen, Yua, Dingwen, Zhang
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916833168195584
author Pan, Liao
Feng, Yang
Wenhui, Zhao
Jinwen, Yua
Dingwen, Zhang
author_facet Pan, Liao
Feng, Yang
Wenhui, Zhao
Jinwen, Yua
Dingwen, Zhang
contents Decoder-only methods, such as GPT, have demonstrated superior performance in many areas compared to traditional encoder-decoder structure transformer methods. Over the years, end-to-end methods based on the traditional transformer structure, like MOTR, have achieved remarkable performance in multi-object tracking. However,the substantial computational resource consumption of these methods, coupled with the optimization challenges posed by dynamic data, results in less favorable inference speeds and training times. To address the aforementioned issues, this paper optimized the network architecture and proposed an effective training strategy to mitigate the problem of prolonged training times, thereby developing DecoderTracker, a novel end-to-end tracking method. Subsequently, to tackle the optimization challenges arising from dynamic data, this paper introduced DecoderTracker+ by incorporating a Fixed-Size Query Memory and refining certain attention layers. Our methods, without any bells and whistles, outperforms MOTR on multiple benchmarks, \textcolor{black}{featuring a 2 to 3 times faster inference than MOTR}, respectively. The proposed method is implemented in open-source code, accessible at https://github.com/liaopan-lp/MO-YOLO.
format Preprint
id arxiv_https___arxiv_org_abs_2310_17170
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle DecoderTracker: Decoder-Only Method for Multiple-Object Tracking
Pan, Liao
Feng, Yang
Wenhui, Zhao
Jinwen, Yua
Dingwen, Zhang
Computer Vision and Pattern Recognition
Decoder-only methods, such as GPT, have demonstrated superior performance in many areas compared to traditional encoder-decoder structure transformer methods. Over the years, end-to-end methods based on the traditional transformer structure, like MOTR, have achieved remarkable performance in multi-object tracking. However,the substantial computational resource consumption of these methods, coupled with the optimization challenges posed by dynamic data, results in less favorable inference speeds and training times. To address the aforementioned issues, this paper optimized the network architecture and proposed an effective training strategy to mitigate the problem of prolonged training times, thereby developing DecoderTracker, a novel end-to-end tracking method. Subsequently, to tackle the optimization challenges arising from dynamic data, this paper introduced DecoderTracker+ by incorporating a Fixed-Size Query Memory and refining certain attention layers. Our methods, without any bells and whistles, outperforms MOTR on multiple benchmarks, \textcolor{black}{featuring a 2 to 3 times faster inference than MOTR}, respectively. The proposed method is implemented in open-source code, accessible at https://github.com/liaopan-lp/MO-YOLO.
title DecoderTracker: Decoder-Only Method for Multiple-Object Tracking
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2310.17170