Video Self-Distillation for Single-Image Encoders: A Step Toward Physically Plausible Perception

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Simon, Marcel, Kim, Tae-Ho, Yeom, Seul-Ki
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916863793954816
author Simon, Marcel
Kim, Tae-Ho
Yeom, Seul-Ki
author_facet Simon, Marcel
Kim, Tae-Ho
Yeom, Seul-Ki
contents Self-supervised image encoders such as DINO have recently gained significant interest for learning robust visual features without labels. However, most SSL methods train on static images and miss the temporal cues inherent in videos. We introduce a video-distilled single-image encoder trained to predict the next-frame representation from the current frame. This simple objective injects 3D spatial and temporal priors without optical flow or tracking. When pre-training on a single 2-hour video, our approach raises the mean Intersection-over-Union (mIoU) on ADE20K from 35.0 (DoRA) to 36.4 while remaining a drop-in replacement for image-only pipelines. Our results highlight video self-distillation as a lightweight route to geometry-aware perception an essential ingredient for physically plausible world models and Physical AI.
format Preprint
id arxiv_https___arxiv_org_abs_2507_19272
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Video Self-Distillation for Single-Image Encoders: A Step Toward Physically Plausible Perception
Simon, Marcel
Kim, Tae-Ho
Yeom, Seul-Ki
Computer Vision and Pattern Recognition
Self-supervised image encoders such as DINO have recently gained significant interest for learning robust visual features without labels. However, most SSL methods train on static images and miss the temporal cues inherent in videos. We introduce a video-distilled single-image encoder trained to predict the next-frame representation from the current frame. This simple objective injects 3D spatial and temporal priors without optical flow or tracking. When pre-training on a single 2-hour video, our approach raises the mean Intersection-over-Union (mIoU) on ADE20K from 35.0 (DoRA) to 36.4 while remaining a drop-in replacement for image-only pipelines. Our results highlight video self-distillation as a lightweight route to geometry-aware perception an essential ingredient for physically plausible world models and Physical AI.
title Video Self-Distillation for Single-Image Encoders: A Step Toward Physically Plausible Perception
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2507.19272