MSPT: Efficient Large-Scale Physical Modeling via Parallelized Multi-Scale Attention
Fuente:
arXiv
Saved in:
| Main Authors: | Curvo, Pedro M. P., van de Meent, Jan-Willem, Zhdanov, Maksim |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Follow the Mean: Reference-Guided Flow Matching
by: Curvo, Pedro M. P., et al.
Published: (2026)
by: Curvo, Pedro M. P., et al.
Published: (2026)
Erwin: A Tree-based Hierarchical Transformer for Large-scale Physical Systems
by: Zhdanov, Maksim, et al.
Published: (2025)
by: Zhdanov, Maksim, et al.
Published: (2025)
(Sparse) Attention to the Details: Preserving Spectral Fidelity in ML-based Weather Forecasting Models
by: Zhdanov, Maksim, et al.
Published: (2026)
by: Zhdanov, Maksim, et al.
Published: (2026)
Crystalite: A Lightweight Transformer for Efficient Crystal Modeling
by: Veljković, Tin Hadži, et al.
Published: (2026)
by: Veljković, Tin Hadži, et al.
Published: (2026)
Conditional Clifford-Steerable CNNs with Complete Kernel Basis for PDE Modeling
by: Szarvas, Bálint László, et al.
Published: (2025)
by: Szarvas, Bálint László, et al.
Published: (2025)
Electrostatics from Laplacian Eigenbasis for Neural Network Interatomic Potentials
by: Zhdanov, Maksim, et al.
Published: (2025)
by: Zhdanov, Maksim, et al.
Published: (2025)
Exponential Family Variational Flow Matching for Tabular Data Generation
by: Guzmán-Cordero, Andrés, et al.
Published: (2025)
by: Guzmán-Cordero, Andrés, et al.
Published: (2025)
VISA: Variational Inference with Sequential Sample-Average Approximations
by: Zimmermann, Heiko, et al.
Published: (2024)
by: Zimmermann, Heiko, et al.
Published: (2024)
Pawsterior: Variational Flow Matching for Structured Simulation-Based Inference
by: Carrasco-Pollo, Jorge, et al.
Published: (2026)
by: Carrasco-Pollo, Jorge, et al.
Published: (2026)
BSA: Ball Sparse Attention for Large-scale Geometries
by: Brita, Catalin E., et al.
Published: (2025)
by: Brita, Catalin E., et al.
Published: (2025)
CORDS: Continuous Representations of Discrete Structures
by: Veljković, Tin Hadži, et al.
Published: (2026)
by: Veljković, Tin Hadži, et al.
Published: (2026)
Adaptive Mesh-Quantization for Neural PDE Solvers
by: Dool, Winfried van den, et al.
Published: (2025)
by: Dool, Winfried van den, et al.
Published: (2025)
Efficient Parallelization Layouts for Large-Scale Distributed Model Training
by: Hagemann, Johannes, et al.
Published: (2023)
by: Hagemann, Johannes, et al.
Published: (2023)
Identity Curvature Laplace Approximation for Improved Out-of-Distribution Detection
by: Zhdanov, Maksim, et al.
Published: (2023)
by: Zhdanov, Maksim, et al.
Published: (2023)
Variational Flow Matching for Graph Generation
by: Eijkelboom, Floor, et al.
Published: (2024)
by: Eijkelboom, Floor, et al.
Published: (2024)
Protocol Models: Scaling Decentralized Training with Communication-Efficient Model Parallelism
by: Ramasinghe, Sameera, et al.
Published: (2025)
by: Ramasinghe, Sameera, et al.
Published: (2025)
Automated Attention Pattern Discovery at Scale in Large Language Models
by: Katzy, Jonathan, et al.
Published: (2026)
by: Katzy, Jonathan, et al.
Published: (2026)
Inverse Concave-Utility Reinforcement Learning is Inverse Game Theory
by: Çelikok, Mustafa Mert, et al.
Published: (2024)
by: Çelikok, Mustafa Mert, et al.
Published: (2024)
Sample-Efficient "Clustering and Conquer" Procedures for Parallel Large-Scale Ranking and Selection
by: Zhang, Zishi, et al.
Published: (2024)
by: Zhang, Zishi, et al.
Published: (2024)
Temporally Multi-Scale Sparse Self-Attention for Physical Activity Data Imputation
by: Wei, Hui, et al.
Published: (2024)
by: Wei, Hui, et al.
Published: (2024)
AMDP: Asynchronous Multi-Directional Pipeline Parallelism for Large-Scale Models Training
by: Chen, Ling, et al.
Published: (2026)
by: Chen, Ling, et al.
Published: (2026)
Kernel-Gradient Drifting Models
by: Esteban-Casadevall, Maria, et al.
Published: (2026)
by: Esteban-Casadevall, Maria, et al.
Published: (2026)
Entropy Coding of Unordered Data Structures
by: Kunze, Julius, et al.
Published: (2024)
by: Kunze, Julius, et al.
Published: (2024)
Piper: Efficient Large-Scale MoE Training via Resource Modeling and Pipelined Hybrid Parallelism
by: Dash, Sajal, et al.
Published: (2026)
by: Dash, Sajal, et al.
Published: (2026)
Scaling Parallel Sequence Models to Foundation-Scale Vision Encoders
by: Jiang, Yitong, et al.
Published: (2026)
by: Jiang, Yitong, et al.
Published: (2026)
AdS-GNN -- a Conformally Equivariant Graph Neural Network
by: Zhdanov, Maksim, et al.
Published: (2025)
by: Zhdanov, Maksim, et al.
Published: (2025)
Diversifying Deep Ensembles: A Saliency Map Approach for Enhanced OOD Detection, Calibration, and Accuracy
by: Dereka, Stanislav, et al.
Published: (2023)
by: Dereka, Stanislav, et al.
Published: (2023)
The Traitors: Deception and Trust in Multi-Agent Language Model Simulations
by: Curvo, Pedro M. P.
Published: (2025)
by: Curvo, Pedro M. P.
Published: (2025)
Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking
by: Ghadia, Ravi, et al.
Published: (2026)
by: Ghadia, Ravi, et al.
Published: (2026)
Parallel Scaling Law for Language Models
by: Chen, Mouxiang, et al.
Published: (2025)
by: Chen, Mouxiang, et al.
Published: (2025)
Large Scale Multi-Task Bayesian Optimization with Large Language Models
by: Zeng, Yimeng, et al.
Published: (2025)
by: Zeng, Yimeng, et al.
Published: (2025)
MoE Parallel Folding: Heterogeneous Parallelism Mappings for Efficient Large-Scale MoE Model Training with Megatron Core
by: Liu, Dennis, et al.
Published: (2025)
by: Liu, Dennis, et al.
Published: (2025)
Two-dimensional Sparse Parallelism for Large Scale Deep Learning Recommendation Model Training
by: Zhang, Xin, et al.
Published: (2025)
by: Zhang, Xin, et al.
Published: (2025)
Multi-Scale and Multimodal Species Distribution Modeling
by: van Tiel, Nina, et al.
Published: (2024)
by: van Tiel, Nina, et al.
Published: (2024)
Scaling Attention via Feature Sparsity
by: Xie, Yan, et al.
Published: (2026)
by: Xie, Yan, et al.
Published: (2026)
Learnable Multipliers: Freeing the Scale of Language Model Matrix Layers
by: Velikanov, Maksim, et al.
Published: (2026)
by: Velikanov, Maksim, et al.
Published: (2026)
Categorical Flow Maps
by: Roos, Daan, et al.
Published: (2026)
by: Roos, Daan, et al.
Published: (2026)
Riemannian Variational Flow Matching for Material and Protein Design
by: Zaghen, Olga, et al.
Published: (2025)
by: Zaghen, Olga, et al.
Published: (2025)
Clifford-Steerable Convolutional Neural Networks
by: Zhdanov, Maksim, et al.
Published: (2024)
by: Zhdanov, Maksim, et al.
Published: (2024)
Predictive Scaling Laws for Efficient GRPO Training of Large Reasoning Models
by: Nimmaturi, Datta, et al.
Published: (2025)
by: Nimmaturi, Datta, et al.
Published: (2025)
Similar Items
-
Follow the Mean: Reference-Guided Flow Matching
by: Curvo, Pedro M. P., et al.
Published: (2026) -
Erwin: A Tree-based Hierarchical Transformer for Large-scale Physical Systems
by: Zhdanov, Maksim, et al.
Published: (2025) -
(Sparse) Attention to the Details: Preserving Spectral Fidelity in ML-based Weather Forecasting Models
by: Zhdanov, Maksim, et al.
Published: (2026) -
Crystalite: A Lightweight Transformer for Efficient Crystal Modeling
by: Veljković, Tin Hadži, et al.
Published: (2026) -
Conditional Clifford-Steerable CNNs with Complete Kernel Basis for PDE Modeling
by: Szarvas, Bálint László, et al.
Published: (2025)