VVitCutLER: Towards Unsupervised Object Detection and Segmentation in Videos

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lu, Zhijing, Hashmi, Khurram Azeem, Stricker, Didier, Afzal, Muhammad Zeshan
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914575709896704
author Lu, Zhijing
Hashmi, Khurram Azeem
Stricker, Didier
Afzal, Muhammad Zeshan
author_facet Lu, Zhijing
Hashmi, Khurram Azeem
Stricker, Didier
Afzal, Muhammad Zeshan
contents Unsupervised pixel-level video understanding remains challenging in real-world scenarios, where motion blur, occlusion, and fast object dynamics often cause temporal drift and flickering pseudo-labels.We propose VVitCutLER, an unsupervised framework for video object detection and instance segmentation, which improves the quality of pseudo-labels through temporal consistency. Our core contribution is VitCut, a temporarily stable pseudo-label generator that reduces error accumulation during field degradation through cross-frame region consistency. Meanwhile, VitCut uses a distillation decoder to achieve effective instance mask prediction. Then, based on VitCut, VVitCutLER further integrates cross-frame feature aggregation to enhance video-level robustness. Extensive experiments on standard video benchmarks demonstrate that VVitCutLER significantly improves detection and segmentation performance while reducing temporal instability. These results highlight the importance of temporally consistent supervision for robust pixel-level video understanding.
format Preprint
id arxiv_https___arxiv_org_abs_2605_17584
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle VVitCutLER: Towards Unsupervised Object Detection and Segmentation in Videos
Lu, Zhijing
Hashmi, Khurram Azeem
Stricker, Didier
Afzal, Muhammad Zeshan
Computer Vision and Pattern Recognition
Unsupervised pixel-level video understanding remains challenging in real-world scenarios, where motion blur, occlusion, and fast object dynamics often cause temporal drift and flickering pseudo-labels.We propose VVitCutLER, an unsupervised framework for video object detection and instance segmentation, which improves the quality of pseudo-labels through temporal consistency. Our core contribution is VitCut, a temporarily stable pseudo-label generator that reduces error accumulation during field degradation through cross-frame region consistency. Meanwhile, VitCut uses a distillation decoder to achieve effective instance mask prediction. Then, based on VitCut, VVitCutLER further integrates cross-frame feature aggregation to enhance video-level robustness. Extensive experiments on standard video benchmarks demonstrate that VVitCutLER significantly improves detection and segmentation performance while reducing temporal instability. These results highlight the importance of temporally consistent supervision for robust pixel-level video understanding.
title VVitCutLER: Towards Unsupervised Object Detection and Segmentation in Videos
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.17584