Algorithmic Task Capture, Computational Complexity, and Inductive Bias of Infinite Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | Davidovich, Orit, Ringel, Zohar |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Towards Understanding Inductive Bias in Transformers: A View From Infinity
by: Lavie, Itay, et al.
Published: (2024)
by: Lavie, Itay, et al.
Published: (2024)
Demystifying Spectral Bias on Real-World Data
by: Lavie, Itay, et al.
Published: (2024)
by: Lavie, Itay, et al.
Published: (2024)
Grokking as a First Order Phase Transition in Two Layer Networks
by: Rubin, Noa, et al.
Published: (2023)
by: Rubin, Noa, et al.
Published: (2023)
From Kernels to Features: A Multi-Scale Adaptive Theory of Feature Learning
by: Rubin, Noa, et al.
Published: (2025)
by: Rubin, Noa, et al.
Published: (2025)
Applications of Statistical Field Theory in Deep Learning
by: Ringel, Zohar, et al.
Published: (2025)
by: Ringel, Zohar, et al.
Published: (2025)
Renormalization group for deep neural networks: Universality of learning and scaling laws
by: Coppola, Gorka Peraza, et al.
Published: (2025)
by: Coppola, Gorka Peraza, et al.
Published: (2025)
Infinite Limits of Multi-head Transformer Dynamics
by: Bordelon, Blake, et al.
Published: (2024)
by: Bordelon, Blake, et al.
Published: (2024)
Lecture notes: From Gaussian processes to feature learning
by: Helias, Moritz, et al.
Published: (2026)
by: Helias, Moritz, et al.
Published: (2026)
Wilsonian Renormalization of Neural Network Gaussian Processes
by: Howard, Jessica N., et al.
Published: (2024)
by: Howard, Jessica N., et al.
Published: (2024)
Transfer Learning in Infinite Width Feature Learning Networks
by: Lauditi, Clarissa, et al.
Published: (2025)
by: Lauditi, Clarissa, et al.
Published: (2025)
Bias in Motion: Theoretical Insights into the Dynamics of Bias in SGD Training
by: Jain, Anchit, et al.
Published: (2024)
by: Jain, Anchit, et al.
Published: (2024)
Critical feature learning in deep neural networks
by: Fischer, Kirsten, et al.
Published: (2024)
by: Fischer, Kirsten, et al.
Published: (2024)
Statistical Physics of Deep Neural Networks: Generalization Capability, Beyond the Infinite Width, and Feature Learning
by: Ariosto, Sebastiano
Published: (2025)
by: Ariosto, Sebastiano
Published: (2025)
Initial Guessing Bias: How Untrained Networks Favor Some Classes
by: Francazi, Emanuele, et al.
Published: (2023)
by: Francazi, Emanuele, et al.
Published: (2023)
Bias-inducing geometries: an exactly solvable data model with fairness implications
by: Mannelli, Stefano Sarao, et al.
Published: (2022)
by: Mannelli, Stefano Sarao, et al.
Published: (2022)
Learning Linear Regression with Low-Rank Tasks in-Context
by: Takanami, Kaito, et al.
Published: (2025)
by: Takanami, Kaito, et al.
Published: (2025)
Exact Learning Dynamics of In-Context Learning in Linear Transformers and Its Application to Non-Linear Transformers
by: Mainali, Nischal, et al.
Published: (2025)
by: Mainali, Nischal, et al.
Published: (2025)
Is Grokking a Computational Glass Relaxation?
by: Zhang, Xiaotian, et al.
Published: (2025)
by: Zhang, Xiaotian, et al.
Published: (2025)
Statistical physics through the lens of real-space mutual information
by: Gökmen, Doruk Efe, et al.
Published: (2021)
by: Gökmen, Doruk Efe, et al.
Published: (2021)
Geometric Dynamics of Signal Propagation Predict Trainability of Transformers
by: Cowsik, Aditya, et al.
Published: (2024)
by: Cowsik, Aditya, et al.
Published: (2024)
Computing frustration and near-monotonicity in deep neural networks
by: Wendin, Joel, et al.
Published: (2025)
by: Wendin, Joel, et al.
Published: (2025)
Small Singular Values Matter: A Random Matrix Analysis of Transformer Models
by: Staats, Max, et al.
Published: (2024)
by: Staats, Max, et al.
Published: (2024)
Computational Thresholds in Multi-Modal Learning via the Spiked Matrix-Tensor Model
by: Tabanelli, Hugo, et al.
Published: (2025)
by: Tabanelli, Hugo, et al.
Published: (2025)
Scaling Laws and Representation Learning in Simple Hierarchical Languages: Transformers vs. Convolutional Architectures
by: Cagnetta, Francesco, et al.
Published: (2025)
by: Cagnetta, Francesco, et al.
Published: (2025)
Rigorous Asymptotics for First-Order Algorithms Through the Dynamical Cavity Method
by: Dandi, Yatin, et al.
Published: (2026)
by: Dandi, Yatin, et al.
Published: (2026)
Compression theory for inhomogeneous systems
by: Gökmen, Doruk Efe, et al.
Published: (2023)
by: Gökmen, Doruk Efe, et al.
Published: (2023)
Generalized Probabilistic Approximate Optimization Algorithm
by: Abdelrahman, Abdelrahman S., et al.
Published: (2025)
by: Abdelrahman, Abdelrahman S., et al.
Published: (2025)
Pattern Expansion of Spin Glasses
by: Shen, Mutian, et al.
Published: (2026)
by: Shen, Mutian, et al.
Published: (2026)
Specialization of softmax attention heads: insights from the high-dimensional single-location model
by: Sagitova, M., et al.
Published: (2026)
by: Sagitova, M., et al.
Published: (2026)
Exact Fixed-Point Constraints in Neural-ODEs with Provable Universality
by: Pacifico, Feliciano Giuseppe, et al.
Published: (2026)
by: Pacifico, Feliciano Giuseppe, et al.
Published: (2026)
Diffusion Operator Geometry of Feedforward Representations
by: Reddy, Kanishka
Published: (2026)
by: Reddy, Kanishka
Published: (2026)
Finite-size scaling of hetero-associative retrieval in continuous-signal-driven Ising spin systems
by: Ladiana, Andrea
Published: (2026)
by: Ladiana, Andrea
Published: (2026)
Emergence of Distortions in High-Dimensional Guided Diffusion Models
by: Ventura, Enrico, et al.
Published: (2026)
by: Ventura, Enrico, et al.
Published: (2026)
A Boundary-Layer Mechanism for One-Third Scaling in Online Softmax Classification
by: Kühn, Marcel, et al.
Published: (2026)
by: Kühn, Marcel, et al.
Published: (2026)
Benchmarking Graph Neural Networks in Solving Hard Constraint Satisfaction Problems
by: Skenderi, Geri, et al.
Published: (2026)
by: Skenderi, Geri, et al.
Published: (2026)
Theory of Optimal Learning Rate Schedules and Scaling Laws for a Random Feature Model
by: Bordelon, Blake, et al.
Published: (2026)
by: Bordelon, Blake, et al.
Published: (2026)
Kernel Renormalization in Bayesian Deep Neural Networks: the Equivalent Wishart Ansatz in the Proportional Regime
by: Baglioni, Paolo, et al.
Published: (2026)
by: Baglioni, Paolo, et al.
Published: (2026)
High-Dimensional Limit of Stochastic Gradient Flow via Dynamical Mean-Field Theory
by: Nishiyama, Sota, et al.
Published: (2026)
by: Nishiyama, Sota, et al.
Published: (2026)
Dynamical Regimes of Multimodal Diffusion Models
by: Albrychiewicz, Emil, et al.
Published: (2026)
by: Albrychiewicz, Emil, et al.
Published: (2026)
Deep Learning as Neural Low-Degree Filtering: A Spectral Theory of Hierarchical Feature Learning
by: Dandi, Yatin, et al.
Published: (2026)
by: Dandi, Yatin, et al.
Published: (2026)
Similar Items
-
Towards Understanding Inductive Bias in Transformers: A View From Infinity
by: Lavie, Itay, et al.
Published: (2024) -
Demystifying Spectral Bias on Real-World Data
by: Lavie, Itay, et al.
Published: (2024) -
Grokking as a First Order Phase Transition in Two Layer Networks
by: Rubin, Noa, et al.
Published: (2023) -
From Kernels to Features: A Multi-Scale Adaptive Theory of Feature Learning
by: Rubin, Noa, et al.
Published: (2025) -
Applications of Statistical Field Theory in Deep Learning
by: Ringel, Zohar, et al.
Published: (2025)