Deriving Transformer Architectures as Implicit Multinomial Regression
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Actor, Jonas A., Gruber, Anthony, Cyr, Eric C. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Gaussian Variational Schemes on Bounded and Unbounded Domains
von: Actor, Jonas A., et al.
Veröffentlicht: (2024)
von: Actor, Jonas A., et al.
Veröffentlicht: (2024)
Multilevel Training for Kolmogorov Arnold Networks
von: Southworth, Ben S., et al.
Veröffentlicht: (2026)
von: Southworth, Ben S., et al.
Veröffentlicht: (2026)
Robust Non-Linear Correlations via Polynomial Regression
von: Giuliani, Luca, et al.
Veröffentlicht: (2025)
von: Giuliani, Luca, et al.
Veröffentlicht: (2025)
Mixture of neural operator experts for learning boundary conditions and model selection
von: Deighan, Dwyer, et al.
Veröffentlicht: (2025)
von: Deighan, Dwyer, et al.
Veröffentlicht: (2025)
Transformers Can Implement Preconditioned Richardson Iteration for In-Context Gaussian Kernel Regression
von: Yan, Mingsong, et al.
Veröffentlicht: (2026)
von: Yan, Mingsong, et al.
Veröffentlicht: (2026)
On the Role of Initialization on the Implicit Bias in Deep Linear Networks
von: Gruber, Oria, et al.
Veröffentlicht: (2024)
von: Gruber, Oria, et al.
Veröffentlicht: (2024)
A Mathematical Explanation of Transformers
von: Tai, Xue-Cheng, et al.
Veröffentlicht: (2025)
von: Tai, Xue-Cheng, et al.
Veröffentlicht: (2025)
Accelerating Matrix Diagonalization through Decision Transformers with Epsilon-Greedy Optimization
von: Bhatta, Kshitij, et al.
Veröffentlicht: (2024)
von: Bhatta, Kshitij, et al.
Veröffentlicht: (2024)
FAST: Factorizable Attention for Speeding up Transformers
von: Gerami, Armin, et al.
Veröffentlicht: (2024)
von: Gerami, Armin, et al.
Veröffentlicht: (2024)
GeoLoRA: Geometric integration for parameter efficient fine-tuning
von: Schotthöfer, Steffen, et al.
Veröffentlicht: (2024)
von: Schotthöfer, Steffen, et al.
Veröffentlicht: (2024)
AlgoFormer: An Efficient Transformer Framework with Algorithmic Structures
von: Gao, Yihang, et al.
Veröffentlicht: (2024)
von: Gao, Yihang, et al.
Veröffentlicht: (2024)
STNet: Spectral Transformation Network for Solving Operator Eigenvalue Problem
von: Wang, Hong, et al.
Veröffentlicht: (2025)
von: Wang, Hong, et al.
Veröffentlicht: (2025)
Principled Approaches for Extending Neural Architectures to Function Spaces for Operator Learning
von: Berner, Julius, et al.
Veröffentlicht: (2025)
von: Berner, Julius, et al.
Veröffentlicht: (2025)
Unisolver: PDE-Conditional Transformers Towards Universal Neural PDE Solvers
von: Zhou, Hang, et al.
Veröffentlicht: (2024)
von: Zhou, Hang, et al.
Veröffentlicht: (2024)
Leveraging KANs for Expedient Training of Multichannel MLPs via Preconditioning and Geometric Refinement
von: Actor, Jonas A., et al.
Veröffentlicht: (2025)
von: Actor, Jonas A., et al.
Veröffentlicht: (2025)
Universal Approximation of Nonlinear Operators and Their Derivatives
von: de Feo, Filippo
Veröffentlicht: (2026)
von: de Feo, Filippo
Veröffentlicht: (2026)
Generalized Orders of Magnitude for Scalable, Parallel, High-Dynamic-Range Computation
von: Heinsen, Franz A., et al.
Veröffentlicht: (2025)
von: Heinsen, Franz A., et al.
Veröffentlicht: (2025)
Learning Explicitly Conditioned Sparsifying Transforms
von: Pătraşcu, Andrei, et al.
Veröffentlicht: (2024)
von: Pătraşcu, Andrei, et al.
Veröffentlicht: (2024)
Towards Faster Matrix Diagonalization with Graph Isomorphism Networks and the AlphaZero Framework
von: Zollicoffer, Geigh, et al.
Veröffentlicht: (2024)
von: Zollicoffer, Geigh, et al.
Veröffentlicht: (2024)
Truncated Matrix Completion - An Empirical Study
von: Naik, Rishhabh, et al.
Veröffentlicht: (2025)
von: Naik, Rishhabh, et al.
Veröffentlicht: (2025)
BWLer: Barycentric Weight Layer Elucidates a Precision-Conditioning Tradeoff for PINNs
von: Liu, Jerry, et al.
Veröffentlicht: (2025)
von: Liu, Jerry, et al.
Veröffentlicht: (2025)
Random weights of DNNs and emergence of fixed points
von: Berlyand, L., et al.
Veröffentlicht: (2025)
von: Berlyand, L., et al.
Veröffentlicht: (2025)
GRAFT: Gradient-Aware Fast MaxVol Technique for Dynamic Data Sampling
von: Jha, Ashish, et al.
Veröffentlicht: (2025)
von: Jha, Ashish, et al.
Veröffentlicht: (2025)
KAN-GCN: Combining Kolmogorov-Arnold Network with Graph Convolution Network for an Accurate Ice Sheet Emulator
von: Liu, Zesheng, et al.
Veröffentlicht: (2025)
von: Liu, Zesheng, et al.
Veröffentlicht: (2025)
Market-Driven Subset Selection for Budgeted Training
von: Jha, Ashish, et al.
Veröffentlicht: (2025)
von: Jha, Ashish, et al.
Veröffentlicht: (2025)
Beyond Loss Guidance: Using PDE Residuals as Spectral Attention in Diffusion Neural Operators
von: Sawhney, Medha, et al.
Veröffentlicht: (2025)
von: Sawhney, Medha, et al.
Veröffentlicht: (2025)
Quasi-Random Physics-informed Neural Networks
von: Yu, Tianchi, et al.
Veröffentlicht: (2025)
von: Yu, Tianchi, et al.
Veröffentlicht: (2025)
Mixed precision accumulation for neural network inference guided by componentwise forward error analysis
von: Arar, El-Mehdi El, et al.
Veröffentlicht: (2025)
von: Arar, El-Mehdi El, et al.
Veröffentlicht: (2025)
Branching Strategies Based on Subgraph GNNs: A Study on Theoretical Promise versus Practical Reality
von: Zhou, Junru, et al.
Veröffentlicht: (2025)
von: Zhou, Junru, et al.
Veröffentlicht: (2025)
Unveiling the Power of Multiple Gossip Steps: A Stability-Based Generalization Analysis in Decentralized Training
von: Li, Qinglun, et al.
Veröffentlicht: (2025)
von: Li, Qinglun, et al.
Veröffentlicht: (2025)
Solving PDEs With Deep Neural Nets under General Boundary Conditions
von: Zhang, Chenggong
Veröffentlicht: (2025)
von: Zhang, Chenggong
Veröffentlicht: (2025)
On Uniform Weighted Deep Polynomial approximation
von: Yeon, Kingsley, et al.
Veröffentlicht: (2025)
von: Yeon, Kingsley, et al.
Veröffentlicht: (2025)
DeepContour: A Hybrid Deep Learning Framework for Accelerating Generalized Eigenvalue Problem Solving via Efficient Contour Design
von: Chen, Yeqiu, et al.
Veröffentlicht: (2025)
von: Chen, Yeqiu, et al.
Veröffentlicht: (2025)
Guided Diffusion Sampling on Function Spaces with Applications to PDEs
von: Yao, Jiachen, et al.
Veröffentlicht: (2025)
von: Yao, Jiachen, et al.
Veröffentlicht: (2025)
Monte Carlo-Type Neural Operator for Differential Equations
von: Choutri, Salah Eddine, et al.
Veröffentlicht: (2025)
von: Choutri, Salah Eddine, et al.
Veröffentlicht: (2025)
Accelerating Eigenvalue Dataset Generation via Chebyshev Subspace Filter
von: Wang, Hong, et al.
Veröffentlicht: (2025)
von: Wang, Hong, et al.
Veröffentlicht: (2025)
Defining Foundation Models for Computational Science: A Call for Clarity and Rigor
von: Choi, Youngsoo, et al.
Veröffentlicht: (2025)
von: Choi, Youngsoo, et al.
Veröffentlicht: (2025)
CFO: Learning Continuous-Time PDE Dynamics via Flow-Matched Neural Operators
von: Hou, Xianglong, et al.
Veröffentlicht: (2025)
von: Hou, Xianglong, et al.
Veröffentlicht: (2025)
ELM-DeepONets: Backpropagation-Free Training of Deep Operator Networks via Extreme Learning Machines
von: Son, Hwijae
Veröffentlicht: (2025)
von: Son, Hwijae
Veröffentlicht: (2025)
Accelerated Gradient-based Design Optimization Via Differentiable Physics-Informed Neural Operator: A Composites Autoclave Processing Case Study
von: Patel, Janak M., et al.
Veröffentlicht: (2025)
von: Patel, Janak M., et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Gaussian Variational Schemes on Bounded and Unbounded Domains
von: Actor, Jonas A., et al.
Veröffentlicht: (2024) -
Multilevel Training for Kolmogorov Arnold Networks
von: Southworth, Ben S., et al.
Veröffentlicht: (2026) -
Robust Non-Linear Correlations via Polynomial Regression
von: Giuliani, Luca, et al.
Veröffentlicht: (2025) -
Mixture of neural operator experts for learning boundary conditions and model selection
von: Deighan, Dwyer, et al.
Veröffentlicht: (2025) -
Transformers Can Implement Preconditioned Richardson Iteration for In-Context Gaussian Kernel Regression
von: Yan, Mingsong, et al.
Veröffentlicht: (2026)