Video Prediction of Dynamic Physical Simulations With Pixel-Space Spatiotemporal Transformers

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Slack, Dean L, Hudson, G Thomas, Winterbottom, Thomas, Moubayed, Noura Al
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914109600038912
author Slack, Dean L
Hudson, G Thomas
Winterbottom, Thomas
Moubayed, Noura Al
author_facet Slack, Dean L
Hudson, G Thomas
Winterbottom, Thomas
Moubayed, Noura Al
contents Inspired by the performance and scalability of autoregressive large language models (LLMs), transformer-based models have seen recent success in the visual domain. This study investigates a transformer adaptation for video prediction with a simple end-to-end approach, comparing various spatiotemporal self-attention layouts. Focusing on causal modeling of physical simulations over time; a common shortcoming of existing video-generative approaches, we attempt to isolate spatiotemporal reasoning via physical object tracking metrics and unsupervised training on physical simulation datasets. We introduce a simple yet effective pure transformer model for autoregressive video prediction, utilizing continuous pixel-space representations for video prediction. Without the need for complex training strategies or latent feature-learning components, our approach significantly extends the time horizon for physically accurate predictions by up to 50% when compared with existing latent-space approaches, while maintaining comparable performance on common video quality metrics. In addition, we conduct interpretability experiments to identify network regions that encode information useful to perform accurate estimations of PDE simulation parameters via probing models, and find that this generalizes to the estimation of out-of-distribution simulation parameters. This work serves as a platform for further attention-based spatiotemporal modeling of videos via a simple, parameter efficient, and interpretable approach.
format Preprint
id arxiv_https___arxiv_org_abs_2510_20807
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Video Prediction of Dynamic Physical Simulations With Pixel-Space Spatiotemporal Transformers
Slack, Dean L
Hudson, G Thomas
Winterbottom, Thomas
Moubayed, Noura Al
Computer Vision and Pattern Recognition
Machine Learning
Inspired by the performance and scalability of autoregressive large language models (LLMs), transformer-based models have seen recent success in the visual domain. This study investigates a transformer adaptation for video prediction with a simple end-to-end approach, comparing various spatiotemporal self-attention layouts. Focusing on causal modeling of physical simulations over time; a common shortcoming of existing video-generative approaches, we attempt to isolate spatiotemporal reasoning via physical object tracking metrics and unsupervised training on physical simulation datasets. We introduce a simple yet effective pure transformer model for autoregressive video prediction, utilizing continuous pixel-space representations for video prediction. Without the need for complex training strategies or latent feature-learning components, our approach significantly extends the time horizon for physically accurate predictions by up to 50% when compared with existing latent-space approaches, while maintaining comparable performance on common video quality metrics. In addition, we conduct interpretability experiments to identify network regions that encode information useful to perform accurate estimations of PDE simulation parameters via probing models, and find that this generalizes to the estimation of out-of-distribution simulation parameters. This work serves as a platform for further attention-based spatiotemporal modeling of videos via a simple, parameter efficient, and interpretable approach.
title Video Prediction of Dynamic Physical Simulations With Pixel-Space Spatiotemporal Transformers
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2510.20807