TAPTR: Tracking Any Point with Transformers as Detection

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Li, Hongyang, Zhang, Hao, Liu, Shilong, Zeng, Zhaoyang, Ren, Tianhe, Li, Feng, Zhang, Lei
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910374729613312
author Li, Hongyang
Zhang, Hao
Liu, Shilong
Zeng, Zhaoyang
Ren, Tianhe
Li, Feng
Zhang, Lei
author_facet Li, Hongyang
Zhang, Hao
Liu, Shilong
Zeng, Zhaoyang
Ren, Tianhe
Li, Feng
Zhang, Lei
contents In this paper, we propose a simple and strong framework for Tracking Any Point with TRansformers (TAPTR). Based on the observation that point tracking bears a great resemblance to object detection and tracking, we borrow designs from DETR-like algorithms to address the task of TAP. In the proposed framework, in each video frame, each tracking point is represented as a point query, which consists of a positional part and a content part. As in DETR, each query (its position and content feature) is naturally updated layer by layer. Its visibility is predicted by its updated content feature. Queries belonging to the same tracking point can exchange information through self-attention along the temporal dimension. As all such operations are well-designed in DETR-like algorithms, the model is conceptually very simple. We also adopt some useful designs such as cost volume from optical flow models and develop simple designs to provide long temporal information while mitigating the feature drifting issue. Our framework demonstrates strong performance with state-of-the-art performance on various TAP datasets with faster inference speed.
format Preprint
id arxiv_https___arxiv_org_abs_2403_13042
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle TAPTR: Tracking Any Point with Transformers as Detection
Li, Hongyang
Zhang, Hao
Liu, Shilong
Zeng, Zhaoyang
Ren, Tianhe
Li, Feng
Zhang, Lei
Computer Vision and Pattern Recognition
Robotics
In this paper, we propose a simple and strong framework for Tracking Any Point with TRansformers (TAPTR). Based on the observation that point tracking bears a great resemblance to object detection and tracking, we borrow designs from DETR-like algorithms to address the task of TAP. In the proposed framework, in each video frame, each tracking point is represented as a point query, which consists of a positional part and a content part. As in DETR, each query (its position and content feature) is naturally updated layer by layer. Its visibility is predicted by its updated content feature. Queries belonging to the same tracking point can exchange information through self-attention along the temporal dimension. As all such operations are well-designed in DETR-like algorithms, the model is conceptually very simple. We also adopt some useful designs such as cost volume from optical flow models and develop simple designs to provide long temporal information while mitigating the feature drifting issue. Our framework demonstrates strong performance with state-of-the-art performance on various TAP datasets with faster inference speed.
title TAPTR: Tracking Any Point with Transformers as Detection
topic Computer Vision and Pattern Recognition
Robotics
url https://arxiv.org/abs/2403.13042