Large Language Models: A Mathematical Formulation
Fuente:
arXiv
Saved in:
| Main Authors: | Baptista, Ricardo, Stuart, Andrew, Tran, Son |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Solving Roughly Forced Nonlinear PDEs via Misspecified Kernel Methods and Neural Networks
by: Baptista, Ricardo, et al.
Published: (2025)
by: Baptista, Ricardo, et al.
Published: (2025)
The Parametric Complexity of Operator Learning
by: Lanthaler, Samuel, et al.
Published: (2023)
by: Lanthaler, Samuel, et al.
Published: (2023)
Nonlocality and Nonlinearity Implies Universality in Operator Learning
by: Lanthaler, Samuel, et al.
Published: (2023)
by: Lanthaler, Samuel, et al.
Published: (2023)
Gradient Flows for Sampling: Mean-Field Models, Gaussian Approximations and Affine Invariance
by: Chen, Yifan, et al.
Published: (2023)
by: Chen, Yifan, et al.
Published: (2023)
Operator Learning: Algorithms and Analysis
by: Kovachki, Nikola B., et al.
Published: (2024)
by: Kovachki, Nikola B., et al.
Published: (2024)
Mathematics of Digital Twins and Transfer Learning for PDE Models
by: Zong, Yifei, et al.
Published: (2025)
by: Zong, Yifei, et al.
Published: (2025)
Numerical Error Analysis of Large Language Models
by: Budzinskiy, Stanislav, et al.
Published: (2025)
by: Budzinskiy, Stanislav, et al.
Published: (2025)
A Kernelizable Primal-Dual Formulation of the Multilinear Singular Value Decomposition
by: Wesel, Frederiek, et al.
Published: (2024)
by: Wesel, Frederiek, et al.
Published: (2024)
A Mathematical Analysis of Neural Operator Behaviors
by: Le, Vu-Anh, et al.
Published: (2024)
by: Le, Vu-Anh, et al.
Published: (2024)
A Mathematical Perspective On Contrastive Learning
by: Baptista, Ricardo, et al.
Published: (2025)
by: Baptista, Ricardo, et al.
Published: (2025)
LAMP: Look-Ahead Mixed-Precision Inference of Large Language Models
by: Budzinskiy, Stanislav, et al.
Published: (2026)
by: Budzinskiy, Stanislav, et al.
Published: (2026)
Continuum Attention for Neural Operators
by: Calvello, Edoardo, et al.
Published: (2024)
by: Calvello, Edoardo, et al.
Published: (2024)
Neural Operator: Learning Maps Between Function Spaces
by: Kovachki, Nikola, et al.
Published: (2021)
by: Kovachki, Nikola, et al.
Published: (2021)
Mathematical analysis of singularities in the diffusion model under the submanifold assumption
by: Lu, Yubin, et al.
Published: (2023)
by: Lu, Yubin, et al.
Published: (2023)
Weak Form Scientific Machine Learning: Test Function Construction for System Identification
by: Tran, April, et al.
Published: (2025)
by: Tran, April, et al.
Published: (2025)
Using Uncertainty Quantification to Characterize and Improve Out-of-Domain Learning for PDEs
by: Mouli, S. Chandra, et al.
Published: (2024)
by: Mouli, S. Chandra, et al.
Published: (2024)
ELM-DeepONets: Backpropagation-Free Training of Deep Operator Networks via Extreme Learning Machines
by: Son, Hwijae
Published: (2025)
by: Son, Hwijae
Published: (2025)
A Mathematical Explanation of Transformers
by: Tai, Xue-Cheng, et al.
Published: (2025)
by: Tai, Xue-Cheng, et al.
Published: (2025)
A Mathematical Guide to Operator Learning
by: Boullé, Nicolas, et al.
Published: (2023)
by: Boullé, Nicolas, et al.
Published: (2023)
Gaussian Measures Conditioned on Nonlinear Observations: Consistency, MAP Estimators, and Simulation
by: Chen, Yifan, et al.
Published: (2024)
by: Chen, Yifan, et al.
Published: (2024)
Numerical Instability and Chaos: Quantifying the Unpredictability of Large Language Models
by: Islam, Chashi Mahiul, et al.
Published: (2026)
by: Islam, Chashi Mahiul, et al.
Published: (2026)
MDD-Thinker: Towards Large Reasoning Models for Major Depressive Disorder Diagnosis
by: Sha, Yuyang, et al.
Published: (2025)
by: Sha, Yuyang, et al.
Published: (2025)
Modeling Large-Scale Walking and Cycling Networks: A Machine Learning Approach Using Mobile Phone and Crowdsourced Data
by: Saberi, Meead, et al.
Published: (2024)
by: Saberi, Meead, et al.
Published: (2024)
Second Order Ensemble Langevin Method for Sampling and Inverse Problems
by: Liu, Ziming, et al.
Published: (2022)
by: Liu, Ziming, et al.
Published: (2022)
Fine-Tune Language Models as Multi-Modal Differential Equation Solvers
by: Yang, Liu, et al.
Published: (2023)
by: Yang, Liu, et al.
Published: (2023)
Pseudo-differential-enhanced physics-informed neural networks
by: Gracyk, Andrew
Published: (2026)
by: Gracyk, Andrew
Published: (2026)
A Machine Learning-Enhanced Hopf-Cole Formulation for Nonlinear Gas Flow in Porous Media
by: Maduru, V. S., et al.
Published: (2026)
by: Maduru, V. S., et al.
Published: (2026)
Operator Learning for Smoothing and Forecasting
by: Calvello, Edoardo, et al.
Published: (2026)
by: Calvello, Edoardo, et al.
Published: (2026)
Sampling via Gradient Flows in the Space of Probability Measures
by: Chen, Yifan, et al.
Published: (2023)
by: Chen, Yifan, et al.
Published: (2023)
Efficient, Multimodal, and Derivative-Free Bayesian Inference With Fisher-Rao Gradient Flows
by: Chen, Yifan, et al.
Published: (2024)
by: Chen, Yifan, et al.
Published: (2024)
Pruning AMR: Efficient Visualization of Implicit Neural Representations via Weight Matrix Analysis
by: Zvonek, Jennifer, et al.
Published: (2025)
by: Zvonek, Jennifer, et al.
Published: (2025)
Gradients of Functions of Large Matrices
by: Krämer, Nicholas, et al.
Published: (2024)
by: Krämer, Nicholas, et al.
Published: (2024)
Mathematical Opportunities in Digital Twins (MATH-DT)
by: Antil, Harbir
Published: (2024)
by: Antil, Harbir
Published: (2024)
Mathematical Modeling of Cancer-Bacterial Therapy: Analysis and Numerical Simulation via Physics-Informed Neural Networks
by: Farkane, Ayoub, et al.
Published: (2026)
by: Farkane, Ayoub, et al.
Published: (2026)
Multi-Level GNN Preconditioner for Solving Large Scale Problems
by: Nastorg, Matthieu, et al.
Published: (2024)
by: Nastorg, Matthieu, et al.
Published: (2024)
Mixture-of-Experts Operator Transformer for Large-Scale PDE Pre-Training
by: Wang, Hong, et al.
Published: (2025)
by: Wang, Hong, et al.
Published: (2025)
Large Data Limits of Laplace Learning for Gaussian Measure Data in Infinite Dimensions
by: Zhong, Zhengang, et al.
Published: (2026)
by: Zhong, Zhengang, et al.
Published: (2026)
Regional climate projections using a deep-learning-based model-ranking and downscaling framework: Application to European climate zones
by: Loganathan, Parthiban, et al.
Published: (2025)
by: Loganathan, Parthiban, et al.
Published: (2025)
Randomized Tensor Ring Decomposition and Its Application to Large-scale Data Reconstruction
by: Yuan, Longhao, et al.
Published: (2019)
by: Yuan, Longhao, et al.
Published: (2019)
DPOT: Auto-Regressive Denoising Operator Transformer for Large-Scale PDE Pre-Training
by: Hao, Zhongkai, et al.
Published: (2024)
by: Hao, Zhongkai, et al.
Published: (2024)
Similar Items
-
Solving Roughly Forced Nonlinear PDEs via Misspecified Kernel Methods and Neural Networks
by: Baptista, Ricardo, et al.
Published: (2025) -
The Parametric Complexity of Operator Learning
by: Lanthaler, Samuel, et al.
Published: (2023) -
Nonlocality and Nonlinearity Implies Universality in Operator Learning
by: Lanthaler, Samuel, et al.
Published: (2023) -
Gradient Flows for Sampling: Mean-Field Models, Gaussian Approximations and Affine Invariance
by: Chen, Yifan, et al.
Published: (2023) -
Operator Learning: Algorithms and Analysis
by: Kovachki, Nikola B., et al.
Published: (2024)