Track-On: Transformer-based Online Point Tracking with Memory

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Aydemir, Görkay, Cai, Xiongyi, Xie, Weidi, Güney, Fatma
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915129490145280
author Aydemir, Görkay
Cai, Xiongyi
Xie, Weidi
Güney, Fatma
author_facet Aydemir, Görkay
Cai, Xiongyi
Xie, Weidi
Güney, Fatma
contents In this paper, we consider the problem of long-term point tracking, which requires consistent identification of points across multiple frames in a video, despite changes in appearance, lighting, perspective, and occlusions. We target online tracking on a frame-by-frame basis, making it suitable for real-world, streaming scenarios. Specifically, we introduce Track-On, a simple transformer-based model designed for online long-term point tracking. Unlike prior methods that depend on full temporal modeling, our model processes video frames causally without access to future frames, leveraging two memory modules -- spatial memory and context memory -- to capture temporal information and maintain reliable point tracking over long time horizons. At inference time, it employs patch classification and refinement to identify correspondences and track points with high accuracy. Through extensive experiments, we demonstrate that Track-On sets a new state-of-the-art for online models and delivers superior or competitive results compared to offline approaches on seven datasets, including the TAP-Vid benchmark. Our method offers a robust and scalable solution for real-time tracking in diverse applications. Project page: https://kuis-ai.github.io/track_on
format Preprint
id arxiv_https___arxiv_org_abs_2501_18487
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Track-On: Transformer-based Online Point Tracking with Memory
Aydemir, Görkay
Cai, Xiongyi
Xie, Weidi
Güney, Fatma
Computer Vision and Pattern Recognition
In this paper, we consider the problem of long-term point tracking, which requires consistent identification of points across multiple frames in a video, despite changes in appearance, lighting, perspective, and occlusions. We target online tracking on a frame-by-frame basis, making it suitable for real-world, streaming scenarios. Specifically, we introduce Track-On, a simple transformer-based model designed for online long-term point tracking. Unlike prior methods that depend on full temporal modeling, our model processes video frames causally without access to future frames, leveraging two memory modules -- spatial memory and context memory -- to capture temporal information and maintain reliable point tracking over long time horizons. At inference time, it employs patch classification and refinement to identify correspondences and track points with high accuracy. Through extensive experiments, we demonstrate that Track-On sets a new state-of-the-art for online models and delivers superior or competitive results compared to offline approaches on seven datasets, including the TAP-Vid benchmark. Our method offers a robust and scalable solution for real-time tracking in diverse applications. Project page: https://kuis-ai.github.io/track_on
title Track-On: Transformer-based Online Point Tracking with Memory
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2501.18487