SAM 2++: Tracking Anything at Any Granularity

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhang, Jiaming, Liang, Cheng, Yang, Yichun, Zeng, Chenkai, Cui, Yutao, Zhang, Xinwen, Zhou, Xin, Ma, Kai, Wu, Gangshan, Wang, Limin
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911691335270400
author Zhang, Jiaming
Liang, Cheng
Yang, Yichun
Zeng, Chenkai
Cui, Yutao
Zhang, Xinwen
Zhou, Xin
Ma, Kai
Wu, Gangshan
Wang, Limin
author_facet Zhang, Jiaming
Liang, Cheng
Yang, Yichun
Zeng, Chenkai
Cui, Yutao
Zhang, Xinwen
Zhou, Xin
Ma, Kai
Wu, Gangshan
Wang, Limin
contents Due to the varying granularity of target states across different tasks, most existing trackers are tailored to a single task, which specificity limits their generalization, preventing them from effectively utilizing multi-task training data and leading to redundancy in both model design and parameters. Although recent unified vision models share partial architectures across tasks, they usually retain task-specific interfaces and overlook the common tracking principle behind different granularities, leaving a gap for truly unified video tracking. To unify video tracking tasks, we present SAM 2++, a unified framework that can handle target states at different granularities, including masks, boxes, and points, through an integrated design of prompt encoding, output decoding, and memory representation. First, to handle different target granularities, we design task-specific prompts that map diverse task inputs into general prompt embeddings, together with a Unified Decoder that produces task results in a common output form without redesigning the overall pipeline. Next, to satisfy memory matching, the core operation of tracking, we introduce a task-adaptive memory mechanism that unifies memory across different granularities while preserving their distinct state semantics, preventing full parameter sharing from causing interference across granularities. Finally, we introduce Tracking-Any-Granularity, the first large and diverse video tracking dataset with rich annotations at three granularities. It is constructed through a customized data engine with phased manual annotation and model-assisted completion, providing a comprehensive resource for training, benchmarking, and analyzing unified tracking models. Comprehensive experiments confirm that SAM 2++ sets a new state of the art across diverse tracking tasks at different granularities, establishing a unified and robust tracking framework.
format Preprint
id arxiv_https___arxiv_org_abs_2510_18822
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SAM 2++: Tracking Anything at Any Granularity
Zhang, Jiaming
Liang, Cheng
Yang, Yichun
Zeng, Chenkai
Cui, Yutao
Zhang, Xinwen
Zhou, Xin
Ma, Kai
Wu, Gangshan
Wang, Limin
Computer Vision and Pattern Recognition
Due to the varying granularity of target states across different tasks, most existing trackers are tailored to a single task, which specificity limits their generalization, preventing them from effectively utilizing multi-task training data and leading to redundancy in both model design and parameters. Although recent unified vision models share partial architectures across tasks, they usually retain task-specific interfaces and overlook the common tracking principle behind different granularities, leaving a gap for truly unified video tracking. To unify video tracking tasks, we present SAM 2++, a unified framework that can handle target states at different granularities, including masks, boxes, and points, through an integrated design of prompt encoding, output decoding, and memory representation. First, to handle different target granularities, we design task-specific prompts that map diverse task inputs into general prompt embeddings, together with a Unified Decoder that produces task results in a common output form without redesigning the overall pipeline. Next, to satisfy memory matching, the core operation of tracking, we introduce a task-adaptive memory mechanism that unifies memory across different granularities while preserving their distinct state semantics, preventing full parameter sharing from causing interference across granularities. Finally, we introduce Tracking-Any-Granularity, the first large and diverse video tracking dataset with rich annotations at three granularities. It is constructed through a customized data engine with phased manual annotation and model-assisted completion, providing a comprehensive resource for training, benchmarking, and analyzing unified tracking models. Comprehensive experiments confirm that SAM 2++ sets a new state of the art across diverse tracking tasks at different granularities, establishing a unified and robust tracking framework.
title SAM 2++: Tracking Anything at Any Granularity
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.18822