On the Value of Tokeniser Pretraining in Physics Foundation Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sotoudeh, Hadi, Mukhopadhyay, Payel, Ohana, Ruben, McCabe, Michael, Lawrence, Neil D., Ho, Shirley, Cranmer, Miles
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908879920562176
author Sotoudeh, Hadi
Mukhopadhyay, Payel
Ohana, Ruben
McCabe, Michael
Lawrence, Neil D.
Ho, Shirley
Cranmer, Miles
author_facet Sotoudeh, Hadi
Mukhopadhyay, Payel
Ohana, Ruben
McCabe, Michael
Lawrence, Neil D.
Ho, Shirley
Cranmer, Miles
contents We investigate the impact of tokeniser pretraining on the accuracy and efficiency of physics emulation. Modern high-resolution simulations produce vast volumes of data spanning diverse physical regimes and scales. Training foundation models to learn the dynamics underlying such data enables the modelling of complex multiphysics phenomena, especially in data-limited settings. The emerging class of physics foundation models typically aims to learn two tasks jointly: (i) extracting compact representations of high-resolution spatiotemporal data, and (ii) capturing governing physical dynamics. However, learning both tasks from scratch simultaneously can impede the effectiveness of either process. We show that pretraining the tokeniser with an autoencoding objective prior to training the dynamics model enhances computational efficiency for physics emulation. Notably, the magnitude of this benefit depends on domain alignment: pretraining on the same physical system as the emulation task yields the largest improvements, while pretraining on other systems provides moderate gains. In-domain pretraining reduces VRMSE by 64% after 10,500 training steps compared to training from scratch. To our knowledge, this is the first systematic investigation of tokeniser pretraining for physics foundation models. We further introduce flexible spatiotemporal compression operations that extend causal convolutions to support runtime-adjustable compression ratios, enabling efficient adaptation to diverse downstream tasks. Our findings provide practical guidance for training efficient physics emulators and highlight the importance of strategic pretraining data selection.
format Preprint
id arxiv_https___arxiv_org_abs_2603_05598
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle On the Value of Tokeniser Pretraining in Physics Foundation Models
Sotoudeh, Hadi
Mukhopadhyay, Payel
Ohana, Ruben
McCabe, Michael
Lawrence, Neil D.
Ho, Shirley
Cranmer, Miles
Machine Learning
Instrumentation and Methods for Astrophysics
Artificial Intelligence
Computational Physics
68T07 (Primary) 68T05, 65M99, 65Z05, 62P35, 35Q35, 35Q31 (Secondary)
I.2.6; I.6.5; I.6.4; I.2.10; G.1.8; F.2.1; I.5.1
We investigate the impact of tokeniser pretraining on the accuracy and efficiency of physics emulation. Modern high-resolution simulations produce vast volumes of data spanning diverse physical regimes and scales. Training foundation models to learn the dynamics underlying such data enables the modelling of complex multiphysics phenomena, especially in data-limited settings. The emerging class of physics foundation models typically aims to learn two tasks jointly: (i) extracting compact representations of high-resolution spatiotemporal data, and (ii) capturing governing physical dynamics. However, learning both tasks from scratch simultaneously can impede the effectiveness of either process. We show that pretraining the tokeniser with an autoencoding objective prior to training the dynamics model enhances computational efficiency for physics emulation. Notably, the magnitude of this benefit depends on domain alignment: pretraining on the same physical system as the emulation task yields the largest improvements, while pretraining on other systems provides moderate gains. In-domain pretraining reduces VRMSE by 64% after 10,500 training steps compared to training from scratch. To our knowledge, this is the first systematic investigation of tokeniser pretraining for physics foundation models. We further introduce flexible spatiotemporal compression operations that extend causal convolutions to support runtime-adjustable compression ratios, enabling efficient adaptation to diverse downstream tasks. Our findings provide practical guidance for training efficient physics emulators and highlight the importance of strategic pretraining data selection.
title On the Value of Tokeniser Pretraining in Physics Foundation Models
topic Machine Learning
Instrumentation and Methods for Astrophysics
Artificial Intelligence
Computational Physics
68T07 (Primary) 68T05, 65M99, 65Z05, 62P35, 35Q35, 35Q31 (Secondary)
I.2.6; I.6.5; I.6.4; I.2.10; G.1.8; F.2.1; I.5.1
url https://arxiv.org/abs/2603.05598