FRAME: Pre-Training Video Feature Representations via Anticipation and Memory

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: TV, Sethuraman, Khosla, Savya, Srinivasakumar, Vignesh, Huang, Jiahui, Oh, Seoung Wug, Jenni, Simon, Hoiem, Derek, Lee, Joon-Young
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913880807047168
author TV, Sethuraman
Khosla, Savya
Srinivasakumar, Vignesh
Huang, Jiahui
Oh, Seoung Wug
Jenni, Simon
Hoiem, Derek
Lee, Joon-Young
author_facet TV, Sethuraman
Khosla, Savya
Srinivasakumar, Vignesh
Huang, Jiahui
Oh, Seoung Wug
Jenni, Simon
Hoiem, Derek
Lee, Joon-Young
contents Dense video prediction tasks, such as object tracking and semantic segmentation, require video encoders that generate temporally consistent, spatially dense features for every frame. However, existing approaches fall short: image encoders like DINO or CLIP lack temporal awareness, while video models such as VideoMAE underperform compared to image encoders on dense prediction tasks. We address this gap with FRAME, a self-supervised video frame encoder tailored for dense video understanding. FRAME learns to predict current and future DINO patch features from past and present RGB frames, leading to spatially precise and temporally coherent representations. To our knowledge, FRAME is the first video encoder to leverage image-based models for dense prediction while outperforming them on tasks requiring fine-grained visual correspondence. As an auxiliary capability, FRAME aligns its class token with CLIP's semantic space, supporting language-driven tasks such as video classification. We evaluate FRAME across six dense prediction tasks on seven datasets, where it consistently outperforms image encoders and existing self-supervised video models. Despite its versatility, FRAME maintains a compact architecture suitable for a range of downstream applications.
format Preprint
id arxiv_https___arxiv_org_abs_2506_05543
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle FRAME: Pre-Training Video Feature Representations via Anticipation and Memory
TV, Sethuraman
Khosla, Savya
Srinivasakumar, Vignesh
Huang, Jiahui
Oh, Seoung Wug
Jenni, Simon
Hoiem, Derek
Lee, Joon-Young
Computer Vision and Pattern Recognition
Dense video prediction tasks, such as object tracking and semantic segmentation, require video encoders that generate temporally consistent, spatially dense features for every frame. However, existing approaches fall short: image encoders like DINO or CLIP lack temporal awareness, while video models such as VideoMAE underperform compared to image encoders on dense prediction tasks. We address this gap with FRAME, a self-supervised video frame encoder tailored for dense video understanding. FRAME learns to predict current and future DINO patch features from past and present RGB frames, leading to spatially precise and temporally coherent representations. To our knowledge, FRAME is the first video encoder to leverage image-based models for dense prediction while outperforming them on tasks requiring fine-grained visual correspondence. As an auxiliary capability, FRAME aligns its class token with CLIP's semantic space, supporting language-driven tasks such as video classification. We evaluate FRAME across six dense prediction tasks on seven datasets, where it consistently outperforms image encoders and existing self-supervised video models. Despite its versatility, FRAME maintains a compact architecture suitable for a range of downstream applications.
title FRAME: Pre-Training Video Feature Representations via Anticipation and Memory
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.05543