Frozen Forecasting: A Unified Evaluation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Walker, Jacob C, Vélez, Pedro, Cabrera, Luisa Polania, Zhou, Guangyao, Ebrahimi, Sayna, Kabra, Rishabh, Doersch, Carl, Ovsjanikov, Maks, Carreira, João, Ginosar, Shiry
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917409765457920
author Walker, Jacob C
Vélez, Pedro
Cabrera, Luisa Polania
Zhou, Guangyao
Ebrahimi, Sayna
Kabra, Rishabh
Doersch, Carl
Ovsjanikov, Maks
Carreira, João
Ginosar, Shiry
author_facet Walker, Jacob C
Vélez, Pedro
Cabrera, Luisa Polania
Zhou, Guangyao
Ebrahimi, Sayna
Kabra, Rishabh
Doersch, Carl
Ovsjanikov, Maks
Carreira, João
Ginosar, Shiry
contents Forecasting future events is a fundamental capability for general-purpose systems that plan or act across different levels of abstraction. Yet, evaluating whether a forecast is "correct" remains challenging due to the inherent uncertainty of the future. We propose a unified evaluation framework for assessing the forecasting capabilities of frozen vision backbones across diverse tasks and abstraction levels. Rather than focusing on single time steps, our framework evaluates entire trajectories and incorporates distributional metrics that better capture the multimodal nature of future outcomes. Given a frozen vision model, we train latent diffusion models to forecast future features directly in its representation space, which are then decoded via lightweight, task-specific readouts. This enables consistent evaluation across a suite of diverse tasks while isolating the forecasting capacity of the backbone itself. We apply our framework to nine diverse vision models, spanning image and video pretraining, contrastive and generative objectives, and with or without language supervision, and evaluate them on four forecasting tasks, from low-level pixel predictions to high-level object motion. We find that forecasting performance strongly correlates with perceptual quality and that the forecasting abilities of video synthesis models are comparable or exceed those pretrained in masking regimes across all levels of abstraction. However, language supervision does not consistently improve forecasting. Notably, video-pretrained models consistently outperform image-based ones.
format Preprint
id arxiv_https___arxiv_org_abs_2507_13942
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Frozen Forecasting: A Unified Evaluation
Walker, Jacob C
Vélez, Pedro
Cabrera, Luisa Polania
Zhou, Guangyao
Ebrahimi, Sayna
Kabra, Rishabh
Doersch, Carl
Ovsjanikov, Maks
Carreira, João
Ginosar, Shiry
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Forecasting future events is a fundamental capability for general-purpose systems that plan or act across different levels of abstraction. Yet, evaluating whether a forecast is "correct" remains challenging due to the inherent uncertainty of the future. We propose a unified evaluation framework for assessing the forecasting capabilities of frozen vision backbones across diverse tasks and abstraction levels. Rather than focusing on single time steps, our framework evaluates entire trajectories and incorporates distributional metrics that better capture the multimodal nature of future outcomes. Given a frozen vision model, we train latent diffusion models to forecast future features directly in its representation space, which are then decoded via lightweight, task-specific readouts. This enables consistent evaluation across a suite of diverse tasks while isolating the forecasting capacity of the backbone itself. We apply our framework to nine diverse vision models, spanning image and video pretraining, contrastive and generative objectives, and with or without language supervision, and evaluate them on four forecasting tasks, from low-level pixel predictions to high-level object motion. We find that forecasting performance strongly correlates with perceptual quality and that the forecasting abilities of video synthesis models are comparable or exceed those pretrained in masking regimes across all levels of abstraction. However, language supervision does not consistently improve forecasting. Notably, video-pretrained models consistently outperform image-based ones.
title Frozen Forecasting: A Unified Evaluation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2507.13942