Mixture-of-Experts Operator Transformer for Large-Scale PDE Pre-Training

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Wang, Hong, Xin, Haiyang, Wang, Jie, Yang, Xuanze, Zha, Fei, Dong, Huanshuo, Jiang, Yan
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866917052462137344
author Wang, Hong
Xin, Haiyang
Wang, Jie
Yang, Xuanze
Zha, Fei
Dong, Huanshuo
Jiang, Yan
author_facet Wang, Hong
Xin, Haiyang
Wang, Jie
Yang, Xuanze
Zha, Fei
Dong, Huanshuo
Jiang, Yan
contents Pre-training has proven effective in addressing data scarcity and performance limitations in solving PDE problems with neural operators. However, challenges remain due to the heterogeneity of PDE datasets in equation types, which leads to high errors in mixed training. Additionally, dense pre-training models that scale parameters by increasing network width or depth incur significant inference costs. To tackle these challenges, we propose a novel Mixture-of-Experts Pre-training Operator Transformer (MoE-POT), a sparse-activated architecture that scales parameters efficiently while controlling inference costs. Specifically, our model adopts a layer-wise router-gating network to dynamically select 4 routed experts from 16 expert networks during inference, enabling the model to focus on equation-specific features. Meanwhile, we also integrate 2 shared experts, aiming to capture common properties of PDE and reduce redundancy among routed experts. The final output is computed as the weighted average of the results from all activated experts. We pre-train models with parameters from 30M to 0.5B on 6 public PDE datasets. Our model with 90M activated parameters achieves up to a 40% reduction in zero-shot error compared with existing models with 120M activated parameters. Additionally, we conduct interpretability analysis, showing that dataset types can be inferred from router-gating network decisions, which validates the rationality and effectiveness of the MoE architecture.
format Preprint
id arxiv_https___arxiv_org_abs_2510_25803
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Mixture-of-Experts Operator Transformer for Large-Scale PDE Pre-Training
Wang, Hong
Xin, Haiyang
Wang, Jie
Yang, Xuanze
Zha, Fei
Dong, Huanshuo
Jiang, Yan
Machine Learning
Numerical Analysis
Pre-training has proven effective in addressing data scarcity and performance limitations in solving PDE problems with neural operators. However, challenges remain due to the heterogeneity of PDE datasets in equation types, which leads to high errors in mixed training. Additionally, dense pre-training models that scale parameters by increasing network width or depth incur significant inference costs. To tackle these challenges, we propose a novel Mixture-of-Experts Pre-training Operator Transformer (MoE-POT), a sparse-activated architecture that scales parameters efficiently while controlling inference costs. Specifically, our model adopts a layer-wise router-gating network to dynamically select 4 routed experts from 16 expert networks during inference, enabling the model to focus on equation-specific features. Meanwhile, we also integrate 2 shared experts, aiming to capture common properties of PDE and reduce redundancy among routed experts. The final output is computed as the weighted average of the results from all activated experts. We pre-train models with parameters from 30M to 0.5B on 6 public PDE datasets. Our model with 90M activated parameters achieves up to a 40% reduction in zero-shot error compared with existing models with 120M activated parameters. Additionally, we conduct interpretability analysis, showing that dataset types can be inferred from router-gating network decisions, which validates the rationality and effectiveness of the MoE architecture.
title Mixture-of-Experts Operator Transformer for Large-Scale PDE Pre-Training
topic Machine Learning
Numerical Analysis
url https://arxiv.org/abs/2510.25803