Saved in:
| Main Authors: | Chen, Lizhang, Li, Jonathan, Liang, Chen, Lao, Ni, Liu, Qiang |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2605.23872 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A PDE-based Explanation of Extreme Numerical Sensitivities and Edge of Stability in Training Neural Networks
by: Sun, Yuxin, et al.
Published: (2022)
by: Sun, Yuxin, et al.
Published: (2022)
$ϕ$-Balancing for Mixture-of-Experts Training
by: Chen, Lizhang, et al.
Published: (2026)
by: Chen, Lizhang, et al.
Published: (2026)
DPOT: Auto-Regressive Denoising Operator Transformer for Large-Scale PDE Pre-Training
by: Hao, Zhongkai, et al.
Published: (2024)
by: Hao, Zhongkai, et al.
Published: (2024)
Variational Matrix-Learning Fourier Networks for Parametric Multiphysics Surrogates
by: Li, Xinyu, et al.
Published: (2026)
by: Li, Xinyu, et al.
Published: (2026)
Mixture-of-Experts Operator Transformer for Large-Scale PDE Pre-Training
by: Wang, Hong, et al.
Published: (2025)
by: Wang, Hong, et al.
Published: (2025)
Positional Knowledge is All You Need: Position-induced Transformer (PiT) for Operator Learning
by: Chen, Junfeng, et al.
Published: (2024)
by: Chen, Junfeng, et al.
Published: (2024)
Time Extrapolation with Graph Convolutional Autoencoder and Tensor Train Decomposition
by: Chen, Yuanhong, et al.
Published: (2025)
by: Chen, Yuanhong, et al.
Published: (2025)
Automatic Differentiation is Essential in Training Neural Networks for Solving Differential Equations
by: Chen, Chuqi, et al.
Published: (2024)
by: Chen, Chuqi, et al.
Published: (2024)
Efficient Transformer-Inspired Variants of Physics-Informed Deep Operator Networks
by: Wei, Zhi-Feng, et al.
Published: (2025)
by: Wei, Zhi-Feng, et al.
Published: (2025)
Training with Fewer Bits: Unlocking Edge LLMs Training with Stochastic Rounding
by: Liu, Taowen, et al.
Published: (2025)
by: Liu, Taowen, et al.
Published: (2025)
Stable Derivative Free Gaussian Mixture Variational Inference for Bayesian Inverse Problems
by: Che, Baojun, et al.
Published: (2025)
by: Che, Baojun, et al.
Published: (2025)
Quantifying Training Difficulty and Accelerating Convergence in Neural Network-Based PDE Solvers
by: Chen, Chuqi, et al.
Published: (2024)
by: Chen, Chuqi, et al.
Published: (2024)
Stabilizing Physics-Informed Consistency Models via Structure-Preserving Training
by: Chang, Che-Chia, et al.
Published: (2026)
by: Chang, Che-Chia, et al.
Published: (2026)
Preconditioning for Accelerated Gradient Descent Optimization and Regularization
by: Ye, Qiang
Published: (2024)
by: Ye, Qiang
Published: (2024)
Diffeomorphism Neural Operator for various domains and parameters of partial differential equations
by: Zhao, Zhiwei, et al.
Published: (2024)
by: Zhao, Zhiwei, et al.
Published: (2024)
IFNSO: Iteration-Free Newton-Schulz Orthogonalization
by: Hu, Chen, et al.
Published: (2026)
by: Hu, Chen, et al.
Published: (2026)
Dual-Domain Deep Learning Method to Accelerate Local Basis Functions Computation for Reservoir Simulation in High-Contrast Porous Media
by: Li, Peiqi, et al.
Published: (2025)
by: Li, Peiqi, et al.
Published: (2025)
Beyond Muon: MUD (MomentUm Decorrelation) for Faster Transformer Training
by: Southworth, Ben S., et al.
Published: (2026)
by: Southworth, Ben S., et al.
Published: (2026)
BCAT: A Block Causal Transformer for PDE Foundation Models for Fluid Dynamics
by: Liu, Yuxuan, et al.
Published: (2025)
by: Liu, Yuxuan, et al.
Published: (2025)
Spectral Estimation with Free Decompression
by: Ameli, Siavash, et al.
Published: (2025)
by: Ameli, Siavash, et al.
Published: (2025)
Free Decompression with Algebraic Spectral Curves
by: Ameli, Siavash, et al.
Published: (2026)
by: Ameli, Siavash, et al.
Published: (2026)
Enhanced BPINN Training Convergence in Solving General and Multi-scale Elliptic PDEs with Noise
by: Hou, Yilong, et al.
Published: (2024)
by: Hou, Yilong, et al.
Published: (2024)
A Mathematical Explanation of Transformers
by: Tai, Xue-Cheng, et al.
Published: (2025)
by: Tai, Xue-Cheng, et al.
Published: (2025)
Training Hamiltonian neural networks without backpropagation
by: Rahma, Atamert, et al.
Published: (2024)
by: Rahma, Atamert, et al.
Published: (2024)
Free-RBF-KAN: Kolmogorov-Arnold Networks with Adaptive Radial Basis Functions for Efficient Function Learning
by: Chiu, Shao-Ting, et al.
Published: (2026)
by: Chiu, Shao-Ting, et al.
Published: (2026)
ELM-DeepONets: Backpropagation-Free Training of Deep Operator Networks via Extreme Learning Machines
by: Son, Hwijae
Published: (2025)
by: Son, Hwijae
Published: (2025)
Multi-Preconditioned LBFGS for Training Finite-Basis PINNs
by: Salvadó-Benasco, Marc, et al.
Published: (2026)
by: Salvadó-Benasco, Marc, et al.
Published: (2026)
Multi-Level Monte Carlo Training of Neural Operators
by: Rowbottom, James, et al.
Published: (2025)
by: Rowbottom, James, et al.
Published: (2025)
Generative Feature Training of Thin 2-Layer Networks
by: Hertrich, Johannes, et al.
Published: (2024)
by: Hertrich, Johannes, et al.
Published: (2024)
Randomized Neural Networks for Integro-Differential Equations with Application to Neutron Transport
by: Dang, Haoning, et al.
Published: (2026)
by: Dang, Haoning, et al.
Published: (2026)
A Computationally Efficient Multidimensional Vision Transformer
by: Ichi, Alaa El, et al.
Published: (2026)
by: Ichi, Alaa El, et al.
Published: (2026)
Power Transform Revisited: Numerically Stable, and Federated
by: Xu, Xuefeng, et al.
Published: (2025)
by: Xu, Xuefeng, et al.
Published: (2025)
Overparameterization of deep ResNet: zero loss and mean-field analysis
by: Ding, Zhiyan, et al.
Published: (2021)
by: Ding, Zhiyan, et al.
Published: (2021)
PirateNets: Physics-informed Deep Learning with Residual Adaptive Networks
by: Wang, Sifan, et al.
Published: (2024)
by: Wang, Sifan, et al.
Published: (2024)
Graph Neural Preconditioners for Iterative Solutions of Sparse Linear Systems
by: Chen, Jie
Published: (2024)
by: Chen, Jie
Published: (2024)
On the Dimension-Free Approximation of Deep Neural Networks for Symmetric Korobov Functions
by: Lu, Yulong, et al.
Published: (2025)
by: Lu, Yulong, et al.
Published: (2025)
Stochastic Dimension Implicit Functional Projections for Exact Integral Conservation in High-Dimensional PINNs
by: Liang, Zhangyong
Published: (2026)
by: Liang, Zhangyong
Published: (2026)
Sharp Convergence Rates for Matching Pursuit
by: Klusowski, Jason M., et al.
Published: (2023)
by: Klusowski, Jason M., et al.
Published: (2023)
Universal Approximation of Operators with Transformers and Neural Integral Operators
by: Zappala, Emanuele, et al.
Published: (2024)
by: Zappala, Emanuele, et al.
Published: (2024)
UFO: A Domain-Unification-Free Operator Framework for Generalized Operator Learning
by: Qiao, Hanli, et al.
Published: (2026)
by: Qiao, Hanli, et al.
Published: (2026)
Similar Items
-
A PDE-based Explanation of Extreme Numerical Sensitivities and Edge of Stability in Training Neural Networks
by: Sun, Yuxin, et al.
Published: (2022) -
$ϕ$-Balancing for Mixture-of-Experts Training
by: Chen, Lizhang, et al.
Published: (2026) -
DPOT: Auto-Regressive Denoising Operator Transformer for Large-Scale PDE Pre-Training
by: Hao, Zhongkai, et al.
Published: (2024) -
Variational Matrix-Learning Fourier Networks for Parametric Multiphysics Surrogates
by: Li, Xinyu, et al.
Published: (2026) -
Mixture-of-Experts Operator Transformer for Large-Scale PDE Pre-Training
by: Wang, Hong, et al.
Published: (2025)