Video Diffusion Models Excel at Tracking Similar-Looking Objects Without Supervision

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhang, Chenshuang, Zhang, Kang, Chung, Joon Son, Kweon, In So, Kim, Junmo, Mao, Chengzhi
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918227722895360
author Zhang, Chenshuang
Zhang, Kang
Chung, Joon Son
Kweon, In So
Kim, Junmo
Mao, Chengzhi
author_facet Zhang, Chenshuang
Zhang, Kang
Chung, Joon Son
Kweon, In So
Kim, Junmo
Mao, Chengzhi
contents Distinguishing visually similar objects by their motion remains a critical challenge in computer vision. Although supervised trackers show promise, contemporary self-supervised trackers struggle when visual cues become ambiguous, limiting their scalability and generalization without extensive labeled data. We find that pre-trained video diffusion models inherently learn motion representations suitable for tracking without task-specific training. This ability arises because their denoising process isolates motion in early, high-noise stages, distinct from later appearance refinement. Capitalizing on this discovery, our self-supervised tracker significantly improves performance in distinguishing visually similar objects, an underexplored failure point for existing methods. Our method achieves up to a 6-point improvement over recent self-supervised approaches on established benchmarks and our newly introduced tests focused on tracking visually similar items. Visualizations confirm that these diffusion-derived motion representations enable robust tracking of even identical objects across challenging viewpoint changes and deformations.
format Preprint
id arxiv_https___arxiv_org_abs_2512_02339
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Video Diffusion Models Excel at Tracking Similar-Looking Objects Without Supervision
Zhang, Chenshuang
Zhang, Kang
Chung, Joon Son
Kweon, In So
Kim, Junmo
Mao, Chengzhi
Computer Vision and Pattern Recognition
Artificial Intelligence
Distinguishing visually similar objects by their motion remains a critical challenge in computer vision. Although supervised trackers show promise, contemporary self-supervised trackers struggle when visual cues become ambiguous, limiting their scalability and generalization without extensive labeled data. We find that pre-trained video diffusion models inherently learn motion representations suitable for tracking without task-specific training. This ability arises because their denoising process isolates motion in early, high-noise stages, distinct from later appearance refinement. Capitalizing on this discovery, our self-supervised tracker significantly improves performance in distinguishing visually similar objects, an underexplored failure point for existing methods. Our method achieves up to a 6-point improvement over recent self-supervised approaches on established benchmarks and our newly introduced tests focused on tracking visually similar items. Visualizations confirm that these diffusion-derived motion representations enable robust tracking of even identical objects across challenging viewpoint changes and deformations.
title Video Diffusion Models Excel at Tracking Similar-Looking Objects Without Supervision
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2512.02339