ARTrackV2: Prompting Autoregressive Tracker Where to Look and How to Describe

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bai, Yifan, Zhao, Zeyang, Gong, Yihong, Wei, Xing
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913233189732352
author Bai, Yifan
Zhao, Zeyang
Gong, Yihong
Wei, Xing
author_facet Bai, Yifan
Zhao, Zeyang
Gong, Yihong
Wei, Xing
contents We present ARTrackV2, which integrates two pivotal aspects of tracking: determining where to look (localization) and how to describe (appearance analysis) the target object across video frames. Building on the foundation of its predecessor, ARTrackV2 extends the concept by introducing a unified generative framework to "read out" object's trajectory and "retell" its appearance in an autoregressive manner. This approach fosters a time-continuous methodology that models the joint evolution of motion and visual features, guided by previous estimates. Furthermore, ARTrackV2 stands out for its efficiency and simplicity, obviating the less efficient intra-frame autoregression and hand-tuned parameters for appearance updates. Despite its simplicity, ARTrackV2 achieves state-of-the-art performance on prevailing benchmark datasets while demonstrating remarkable efficiency improvement. In particular, ARTrackV2 achieves AO score of 79.5\% on GOT-10k, and AUC of 86.1\% on TrackingNet while being $3.6 \times$ faster than ARTrack. The code will be released.
format Preprint
id arxiv_https___arxiv_org_abs_2312_17133
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle ARTrackV2: Prompting Autoregressive Tracker Where to Look and How to Describe
Bai, Yifan
Zhao, Zeyang
Gong, Yihong
Wei, Xing
Computer Vision and Pattern Recognition
We present ARTrackV2, which integrates two pivotal aspects of tracking: determining where to look (localization) and how to describe (appearance analysis) the target object across video frames. Building on the foundation of its predecessor, ARTrackV2 extends the concept by introducing a unified generative framework to "read out" object's trajectory and "retell" its appearance in an autoregressive manner. This approach fosters a time-continuous methodology that models the joint evolution of motion and visual features, guided by previous estimates. Furthermore, ARTrackV2 stands out for its efficiency and simplicity, obviating the less efficient intra-frame autoregression and hand-tuned parameters for appearance updates. Despite its simplicity, ARTrackV2 achieves state-of-the-art performance on prevailing benchmark datasets while demonstrating remarkable efficiency improvement. In particular, ARTrackV2 achieves AO score of 79.5\% on GOT-10k, and AUC of 86.1\% on TrackingNet while being $3.6 \times$ faster than ARTrack. The code will be released.
title ARTrackV2: Prompting Autoregressive Tracker Where to Look and How to Describe
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2312.17133