MoSiC: Optimal-Transport Motion Trajectory for Dense Self-Supervised Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Salehi, Mohammadreza, Venkataramanan, Shashanka, Simion, Ioana, Gavves, Efstratios, Snoek, Cees G. M., Asano, Yuki M
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911049059401728
author Salehi, Mohammadreza
Venkataramanan, Shashanka
Simion, Ioana
Gavves, Efstratios
Snoek, Cees G. M.
Asano, Yuki M
author_facet Salehi, Mohammadreza
Venkataramanan, Shashanka
Simion, Ioana
Gavves, Efstratios
Snoek, Cees G. M.
Asano, Yuki M
contents Dense self-supervised learning has shown great promise for learning pixel- and patch-level representations, but extending it to videos remains challenging due to the complexity of motion dynamics. Existing approaches struggle as they rely on static augmentations that fail under object deformations, occlusions, and camera movement, leading to inconsistent feature learning over time. We propose a motion-guided self-supervised learning framework that clusters dense point tracks to learn spatiotemporally consistent representations. By leveraging an off-the-shelf point tracker, we extract long-range motion trajectories and optimize feature clustering through a momentum-encoder-based optimal transport mechanism. To ensure temporal coherence, we propagate cluster assignments along tracked points, enforcing feature consistency across views despite viewpoint changes. Integrating motion as an implicit supervisory signal, our method learns representations that generalize across frames, improving robustness in dynamic scenes and challenging occlusion scenarios. By initializing from strong image-pretrained models and leveraging video data for training, we improve state-of-the-art by 1% to 6% on six image and video datasets and four evaluation benchmarks. The implementation is publicly available at our GitHub repository: https://github.com/SMSD75/MoSiC/tree/main
format Preprint
id arxiv_https___arxiv_org_abs_2506_08694
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MoSiC: Optimal-Transport Motion Trajectory for Dense Self-Supervised Learning
Salehi, Mohammadreza
Venkataramanan, Shashanka
Simion, Ioana
Gavves, Efstratios
Snoek, Cees G. M.
Asano, Yuki M
Computer Vision and Pattern Recognition
Dense self-supervised learning has shown great promise for learning pixel- and patch-level representations, but extending it to videos remains challenging due to the complexity of motion dynamics. Existing approaches struggle as they rely on static augmentations that fail under object deformations, occlusions, and camera movement, leading to inconsistent feature learning over time. We propose a motion-guided self-supervised learning framework that clusters dense point tracks to learn spatiotemporally consistent representations. By leveraging an off-the-shelf point tracker, we extract long-range motion trajectories and optimize feature clustering through a momentum-encoder-based optimal transport mechanism. To ensure temporal coherence, we propagate cluster assignments along tracked points, enforcing feature consistency across views despite viewpoint changes. Integrating motion as an implicit supervisory signal, our method learns representations that generalize across frames, improving robustness in dynamic scenes and challenging occlusion scenarios. By initializing from strong image-pretrained models and leveraging video data for training, we improve state-of-the-art by 1% to 6% on six image and video datasets and four evaluation benchmarks. The implementation is publicly available at our GitHub repository: https://github.com/SMSD75/MoSiC/tree/main
title MoSiC: Optimal-Transport Motion Trajectory for Dense Self-Supervised Learning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.08694