MSPT: Efficient Large-Scale Physical Modeling via Parallelized Multi-Scale Attention

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Curvo, Pedro M. P., van de Meent, Jan-Willem, Zhdanov, Maksim
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908872847917056
author Curvo, Pedro M. P.
van de Meent, Jan-Willem
Zhdanov, Maksim
author_facet Curvo, Pedro M. P.
van de Meent, Jan-Willem
Zhdanov, Maksim
contents A key scalability challenge in neural solvers for industrial-scale physics simulations is efficiently capturing both fine-grained local interactions and long-range global dependencies across millions of spatial elements. We introduce the Multi-Scale Patch Transformer (MSPT), an architecture that combines local point attention within patches with global attention to coarse patch-level representations. To partition the input domain into spatially-coherent patches, we employ ball trees, which handle irregular geometries efficiently. This dual-scale design enables MSPT to scale to millions of points on a single GPU. We validate our method on standard PDE benchmarks (elasticity, plasticity, fluid dynamics, porous flow) and large-scale aerodynamic datasets (ShapeNet-Car, Ahmed-ML), achieving state-of-the-art accuracy with substantially lower memory footprint and computational cost.
format Preprint
id arxiv_https___arxiv_org_abs_2512_01738
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MSPT: Efficient Large-Scale Physical Modeling via Parallelized Multi-Scale Attention
Curvo, Pedro M. P.
van de Meent, Jan-Willem
Zhdanov, Maksim
Machine Learning
A key scalability challenge in neural solvers for industrial-scale physics simulations is efficiently capturing both fine-grained local interactions and long-range global dependencies across millions of spatial elements. We introduce the Multi-Scale Patch Transformer (MSPT), an architecture that combines local point attention within patches with global attention to coarse patch-level representations. To partition the input domain into spatially-coherent patches, we employ ball trees, which handle irregular geometries efficiently. This dual-scale design enables MSPT to scale to millions of points on a single GPU. We validate our method on standard PDE benchmarks (elasticity, plasticity, fluid dynamics, porous flow) and large-scale aerodynamic datasets (ShapeNet-Car, Ahmed-ML), achieving state-of-the-art accuracy with substantially lower memory footprint and computational cost.
title MSPT: Efficient Large-Scale Physical Modeling via Parallelized Multi-Scale Attention
topic Machine Learning
url https://arxiv.org/abs/2512.01738