Video Prediction of Dynamic Physical Simulations With Pixel-Space Spatiotemporal Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | Slack, Dean L, Hudson, G Thomas, Winterbottom, Thomas, Moubayed, Noura Al |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Everything is a Video: Unifying Modalities through Next-Frame Prediction
by: Hudson, G. Thomas, et al.
Published: (2024)
by: Hudson, G. Thomas, et al.
Published: (2024)
The Power of Next-Frame Prediction for Learning Physical Laws
by: Winterbottom, Thomas, et al.
Published: (2024)
by: Winterbottom, Thomas, et al.
Published: (2024)
Controllable Image Generation with Composed Parallel Token Prediction
by: Stirling, Jamie, et al.
Published: (2024)
by: Stirling, Jamie, et al.
Published: (2024)
Investigating Permutation-Invariant Discrete Representation Learning for Spatially Aligned Images
by: Stirling, Jamie S. J., et al.
Published: (2026)
by: Stirling, Jamie S. J., et al.
Published: (2026)
Disentangling Racial Phenotypes: Fine-Grained Control of Race-related Facial Phenotype Characteristics
by: Yucer, Seyma, et al.
Published: (2024)
by: Yucer, Seyma, et al.
Published: (2024)
Pixel Sentence Representation Learning
by: Xiao, Chenghao, et al.
Published: (2024)
by: Xiao, Chenghao, et al.
Published: (2024)
AttenCraft: Attention-guided Disentanglement of Multiple Concepts for Text-to-Image Customization
by: Shentu, Junjie, et al.
Published: (2024)
by: Shentu, Junjie, et al.
Published: (2024)
Textual Localization: Decomposing Multi-concept Images for Subject-Driven Text-to-Image Generation
by: Shentu, Junjie, et al.
Published: (2024)
by: Shentu, Junjie, et al.
Published: (2024)
Physics-Driven Spatiotemporal Modeling for AI-Generated Video Detection
by: Zhang, Shuhai, et al.
Published: (2025)
by: Zhang, Shuhai, et al.
Published: (2025)
Early Detection and Reduction of Memorisation for Domain Adaptation and Instruction Tuning
by: Slack, Dean L., et al.
Published: (2025)
by: Slack, Dean L., et al.
Published: (2025)
Scalable High-Resolution Pixel-Space Image Synthesis with Hourglass Diffusion Transformers
by: Crowson, Katherine, et al.
Published: (2024)
by: Crowson, Katherine, et al.
Published: (2024)
Minerva-Ego: Spatiotemporal Hints for Egocentric Video Understanding
by: Nagrani, Arsha, et al.
Published: (2026)
by: Nagrani, Arsha, et al.
Published: (2026)
Latent Forcing: Reordering the Diffusion Trajectory for Pixel-Space Image Generation
by: Baade, Alan, et al.
Published: (2026)
by: Baade, Alan, et al.
Published: (2026)
Improving Out-of-Domain Robustness with Targeted Augmentation in Frequency and Pixel Spaces
by: Wang, Ruoqi, et al.
Published: (2025)
by: Wang, Ruoqi, et al.
Published: (2025)
FREPix: Frequency-Heterogeneous Flow Matching for Pixel-Space Image Generation
by: Lin, Mingfeng, et al.
Published: (2026)
by: Lin, Mingfeng, et al.
Published: (2026)
CoordFlow: Coordinate Flow for Pixel-wise Neural Video Representation
by: Silver, Daniel, et al.
Published: (2025)
by: Silver, Daniel, et al.
Published: (2025)
Representation Learning for Spatiotemporal Physical Systems
by: Qu, Helen, et al.
Published: (2026)
by: Qu, Helen, et al.
Published: (2026)
Model Predictive Simulation Using Structured Graphical Models and Transformers
by: Lou, Xinghua, et al.
Published: (2024)
by: Lou, Xinghua, et al.
Published: (2024)
An Image is Worth More Than 16x16 Patches: Exploring Transformers on Individual Pixels
by: Nguyen, Duy-Kien, et al.
Published: (2024)
by: Nguyen, Duy-Kien, et al.
Published: (2024)
EVEREST: Efficient Masked Video Autoencoder by Removing Redundant Spatiotemporal Tokens
by: Hwang, Sunil, et al.
Published: (2022)
by: Hwang, Sunil, et al.
Published: (2022)
Edify Image: High-Quality Image Generation with Pixel Space Laplacian Diffusion Models
by: NVIDIA, et al.
Published: (2024)
by: NVIDIA, et al.
Published: (2024)
UCTB: An Urban Computing Tool Box for Building Spatiotemporal Prediction Services
by: Fang, Jiangyi, et al.
Published: (2023)
by: Fang, Jiangyi, et al.
Published: (2023)
Mixture of Distributions Matters: Dynamic Sparse Attention for Efficient Video Diffusion Transformers
by: Liu, Yuxi, et al.
Published: (2026)
by: Liu, Yuxi, et al.
Published: (2026)
Featurising Pixels from Dynamic 3D Scenes with Linear In-Context Learners
by: Araslanov, Nikita, et al.
Published: (2026)
by: Araslanov, Nikita, et al.
Published: (2026)
From Pixels to Perception: Interpretable Predictions via Instance-wise Grouped Feature Selection
by: Vandenhirtz, Moritz, et al.
Published: (2025)
by: Vandenhirtz, Moritz, et al.
Published: (2025)
Inferring Dynamic Physical Properties from Video Foundation Models
by: Zhan, Guanqi, et al.
Published: (2025)
by: Zhan, Guanqi, et al.
Published: (2025)
Spatiotemporal Tile-based Attention-guided LSTMs for Traffic Video Prediction
by: Nguyen, Tu
Published: (2019)
by: Nguyen, Tu
Published: (2019)
Pixel-Space Post-Training of Latent Diffusion Models
by: Zhang, Christina, et al.
Published: (2024)
by: Zhang, Christina, et al.
Published: (2024)
DDLP: Unsupervised Object-Centric Video Prediction with Deep Dynamic Latent Particles
by: Daniel, Tal, et al.
Published: (2023)
by: Daniel, Tal, et al.
Published: (2023)
What's in a Latent? Leveraging Diffusion Latent Space for Domain Generalization
by: Thomas, Xavier, et al.
Published: (2025)
by: Thomas, Xavier, et al.
Published: (2025)
LAPA: Log-Domain Prediction-Driven Dynamic Sparsity Accelerator for Transformer Model
by: Wang, Huizheng, et al.
Published: (2025)
by: Wang, Huizheng, et al.
Published: (2025)
From Pixels to BFS: High Maze Accuracy Does Not Imply Visual Planning
by: Salgado, Alberto G. Rodriguez
Published: (2026)
by: Salgado, Alberto G. Rodriguez
Published: (2026)
Rethinking Generative Image Pretraining: How Far Are We From Scaling Up Next-Pixel Prediction?
by: Yan, Xinchen, et al.
Published: (2025)
by: Yan, Xinchen, et al.
Published: (2025)
Context-Aware Zero-Shot Anomaly Detection in Surveillance Using Contrastive and Predictive Spatiotemporal Modeling
by: Khan, Md. Rashid Shahriar, et al.
Published: (2025)
by: Khan, Md. Rashid Shahriar, et al.
Published: (2025)
Parallelized Spatiotemporal Binding
by: Singh, Gautam, et al.
Published: (2024)
by: Singh, Gautam, et al.
Published: (2024)
Associative Memories in the Feature Space
by: Salvatori, Tommaso, et al.
Published: (2024)
by: Salvatori, Tommaso, et al.
Published: (2024)
VideoOrion: Tokenizing Object Dynamics in Videos
by: Feng, Yicheng, et al.
Published: (2024)
by: Feng, Yicheng, et al.
Published: (2024)
Video-Driven Graph Network-Based Simulators
by: Szewczyk, Franciszek, et al.
Published: (2024)
by: Szewczyk, Franciszek, et al.
Published: (2024)
SSL-AD: Spatiotemporal Self-Supervised Learning for Generalizability and Adaptability Across Alzheimer's Prediction Tasks and Datasets
by: Kaczmarek, Emily, et al.
Published: (2025)
by: Kaczmarek, Emily, et al.
Published: (2025)
Block-Recurrent Dynamics in Vision Transformers
by: Jacobs, Mozes, et al.
Published: (2025)
by: Jacobs, Mozes, et al.
Published: (2025)
Similar Items
-
Everything is a Video: Unifying Modalities through Next-Frame Prediction
by: Hudson, G. Thomas, et al.
Published: (2024) -
The Power of Next-Frame Prediction for Learning Physical Laws
by: Winterbottom, Thomas, et al.
Published: (2024) -
Controllable Image Generation with Composed Parallel Token Prediction
by: Stirling, Jamie, et al.
Published: (2024) -
Investigating Permutation-Invariant Discrete Representation Learning for Spatially Aligned Images
by: Stirling, Jamie S. J., et al.
Published: (2026) -
Disentangling Racial Phenotypes: Fine-Grained Control of Race-related Facial Phenotype Characteristics
by: Yucer, Seyma, et al.
Published: (2024)