DINO-Foresight: Looking into the Future with DINO

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Karypidis, Efstathios, Kakogeorgiou, Ioannis, Gidaris, Spyros, Komodakis, Nikos
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911292305965056
author Karypidis, Efstathios
Kakogeorgiou, Ioannis
Gidaris, Spyros
Komodakis, Nikos
author_facet Karypidis, Efstathios
Kakogeorgiou, Ioannis
Gidaris, Spyros
Komodakis, Nikos
contents Predicting future dynamics is crucial for applications like autonomous driving and robotics, where understanding the environment is key. Existing pixel-level methods are computationally expensive and often focus on irrelevant details. To address these challenges, we introduce DINO-Foresight, a novel framework that operates in the semantic feature space of pretrained Vision Foundation Models (VFMs). Our approach trains a masked feature transformer in a self-supervised manner to predict the evolution of VFM features over time. By forecasting these features, we can apply off-the-shelf, task-specific heads for various scene understanding tasks. In this framework, VFM features are treated as a latent space, to which different heads attach to perform specific tasks for future-frame analysis. Extensive experiments show the very strong performance, robustness and scalability of our framework. Project page and code at https://dino-foresight.github.io/ .
format Preprint
id arxiv_https___arxiv_org_abs_2412_11673
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle DINO-Foresight: Looking into the Future with DINO
Karypidis, Efstathios
Kakogeorgiou, Ioannis
Gidaris, Spyros
Komodakis, Nikos
Computer Vision and Pattern Recognition
Predicting future dynamics is crucial for applications like autonomous driving and robotics, where understanding the environment is key. Existing pixel-level methods are computationally expensive and often focus on irrelevant details. To address these challenges, we introduce DINO-Foresight, a novel framework that operates in the semantic feature space of pretrained Vision Foundation Models (VFMs). Our approach trains a masked feature transformer in a self-supervised manner to predict the evolution of VFM features over time. By forecasting these features, we can apply off-the-shelf, task-specific heads for various scene understanding tasks. In this framework, VFM features are treated as a latent space, to which different heads attach to perform specific tasks for future-frame analysis. Extensive experiments show the very strong performance, robustness and scalability of our framework. Project page and code at https://dino-foresight.github.io/ .
title DINO-Foresight: Looking into the Future with DINO
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.11673