The Persian Rug: solving toy models of superposition using large-scale symmetries
Fuente:
arXiv
Saved in:
| Main Authors: | Cowsik, Aditya, Dolev, Kfir, Infanger, Alex |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Towards Distributed Neural Architectures
by: Cowsik, Aditya, et al.
Published: (2025)
by: Cowsik, Aditya, et al.
Published: (2025)
Geometric Dynamics of Signal Propagation Predict Trainability of Transformers
by: Cowsik, Aditya, et al.
Published: (2024)
by: Cowsik, Aditya, et al.
Published: (2024)
A method for quantifying the generalization capabilities of generative models for solving Ising models
by: Ma, Qunlong, et al.
Published: (2024)
by: Ma, Qunlong, et al.
Published: (2024)
On the origin of neural scaling laws: from random graphs to natural language
by: Barkeshli, Maissam, et al.
Published: (2026)
by: Barkeshli, Maissam, et al.
Published: (2026)
Generalization through variance: how noise shapes inductive biases in diffusion models
by: Vastola, John J.
Published: (2025)
by: Vastola, John J.
Published: (2025)
Identifying internal patterns in (1+1)-dimensional directed percolation using neural networks
by: Parkhomenko, Danil, et al.
Published: (2025)
by: Parkhomenko, Danil, et al.
Published: (2025)
More Bang for the Buck: Improving the Inference of Large Language Models at a Fixed Budget using Reset and Discard (ReD)
by: Meir, Sagi, et al.
Published: (2026)
by: Meir, Sagi, et al.
Published: (2026)
KAN: Kolmogorov-Arnold Networks
by: Liu, Ziming, et al.
Published: (2024)
by: Liu, Ziming, et al.
Published: (2024)
How Do Transformers "Do" Physics? Investigating the Simple Harmonic Oscillator
by: Kantamneni, Subhash, et al.
Published: (2024)
by: Kantamneni, Subhash, et al.
Published: (2024)
Smooth Kolmogorov Arnold networks enabling structural knowledge representation
by: Samadi, Moein E., et al.
Published: (2024)
by: Samadi, Moein E., et al.
Published: (2024)
Scaling Laws and Spectra of Shallow Neural Networks in the Feature Learning Regime
by: Defilippis, Leonardo, et al.
Published: (2025)
by: Defilippis, Leonardo, et al.
Published: (2025)
A Scalable Measure of Loss Landscape Curvature for Analyzing the Training Dynamics of LLMs
by: Kalra, Dayal Singh, et al.
Published: (2026)
by: Kalra, Dayal Singh, et al.
Published: (2026)
Grokking vs. Learning: Same Features, Different Encodings
by: Manning-Coe, Dmitry, et al.
Published: (2025)
by: Manning-Coe, Dmitry, et al.
Published: (2025)
Quantifying Hyperparameter Transfer and the Importance of Embedding Layer Learning Rate
by: Kalra, Dayal Singh, et al.
Published: (2026)
by: Kalra, Dayal Singh, et al.
Published: (2026)
A Spin Glass Characterization of Neural Networks
by: Li, Jun
Published: (2025)
by: Li, Jun
Published: (2025)
Applications of Statistical Field Theory in Deep Learning
by: Ringel, Zohar, et al.
Published: (2025)
by: Ringel, Zohar, et al.
Published: (2025)
Spectral Dynamics in Deep Networks: Feature Learning, Outlier Escape, and Learning Rate Transfer
by: Lauditi, Clarissa, et al.
Published: (2026)
by: Lauditi, Clarissa, et al.
Published: (2026)
A Geometric Perspective on the Difficulties of Learning GNN-based SAT Solvers
by: Skenderi, Geri
Published: (2025)
by: Skenderi, Geri
Published: (2025)
Learning Shrinks the Hard Tail: Training-Dependent Inference Scaling in a Solvable Linear Model
by: Levi, Noam
Published: (2026)
by: Levi, Noam
Published: (2026)
Enhancing Noise-Robust Losses for Large-Scale Noisy Data Learning
by: Staats, Max, et al.
Published: (2023)
by: Staats, Max, et al.
Published: (2023)
Representation Learning on a Random Lattice
by: Brill, Aryeh
Published: (2025)
by: Brill, Aryeh
Published: (2025)
Connecting NTK and NNGP: A Unified Theoretical Framework for Wide Neural Network Learning Dynamics
by: Avidan, Yehonatan, et al.
Published: (2023)
by: Avidan, Yehonatan, et al.
Published: (2023)
Parameter Symmetry Potentially Unifies Deep Learning Theory
by: Ziyin, Liu, et al.
Published: (2025)
by: Ziyin, Liu, et al.
Published: (2025)
When resampling/reweighting improves feature learning in imbalanced classification?: A toy-model study
by: Obuchi, Tomoyuki, et al.
Published: (2024)
by: Obuchi, Tomoyuki, et al.
Published: (2024)
Differential learning kinetics govern the transition from memorization to generalization during in-context learning
by: Nguyen, Alex, et al.
Published: (2024)
by: Nguyen, Alex, et al.
Published: (2024)
Nature-Inspired Local Propagation
by: Betti, Alessandro, et al.
Published: (2024)
by: Betti, Alessandro, et al.
Published: (2024)
Predictive Coding Networks and Inference Learning: Tutorial and Survey
by: van Zwol, Björn, et al.
Published: (2024)
by: van Zwol, Björn, et al.
Published: (2024)
Predictive Coding Graphs are a Superset of Feedforward Neural Networks
by: van Zwol, Björn
Published: (2026)
by: van Zwol, Björn
Published: (2026)
Approximation Theory for Neural Networks: Old and New
by: Mukherjee, Soumendu Sundar, et al.
Published: (2026)
by: Mukherjee, Soumendu Sundar, et al.
Published: (2026)
Preisach Attention: A Hysteretic Model of Sequential Memory
by: Frydrych, Piotr
Published: (2026)
by: Frydrych, Piotr
Published: (2026)
Lattice Protein Folding with Variational Annealing
by: Khandoker, Shoummo Ahsan, et al.
Published: (2025)
by: Khandoker, Shoummo Ahsan, et al.
Published: (2025)
Non-equilibrium active noise enhances generative memory in diffusion models
by: Behera, Agnish Kumar, et al.
Published: (2024)
by: Behera, Agnish Kumar, et al.
Published: (2024)
An exactly solvable model for emergence and scaling laws in the multitask sparse parity problem
by: Nam, Yoonsoo, et al.
Published: (2024)
by: Nam, Yoonsoo, et al.
Published: (2024)
Learning and extrapolating scale-invariant processes
by: Alvez-Canepa, Anaclara, et al.
Published: (2026)
by: Alvez-Canepa, Anaclara, et al.
Published: (2026)
Autonomous Agents Coordinating Distributed Discovery Through Emergent Artifact Exchange
by: Wang, Fiona Y., et al.
Published: (2026)
by: Wang, Fiona Y., et al.
Published: (2026)
Sampling at intermediate temperatures is optimal for training large language models in protein structure prediction
by: Ghiringhelli, L., et al.
Published: (2026)
by: Ghiringhelli, L., et al.
Published: (2026)
Critical behavior of dirty parafermionic chains
by: Pandey, Akshat, et al.
Published: (2024)
by: Pandey, Akshat, et al.
Published: (2024)
Dimension-free deterministic equivalents and scaling laws for random feature regression
by: Defilippis, Leonardo, et al.
Published: (2024)
by: Defilippis, Leonardo, et al.
Published: (2024)
Deterministic versus stochastic dynamical classifiers: opposing random adversarial attacks with noise
by: Chicchi, Lorenzo, et al.
Published: (2024)
by: Chicchi, Lorenzo, et al.
Published: (2024)
Statistical Mechanics and Artificial Neural Networks: Principles, Models, and Applications
by: Böttcher, Lucas, et al.
Published: (2024)
by: Böttcher, Lucas, et al.
Published: (2024)
Similar Items
-
Towards Distributed Neural Architectures
by: Cowsik, Aditya, et al.
Published: (2025) -
Geometric Dynamics of Signal Propagation Predict Trainability of Transformers
by: Cowsik, Aditya, et al.
Published: (2024) -
A method for quantifying the generalization capabilities of generative models for solving Ising models
by: Ma, Qunlong, et al.
Published: (2024) -
On the origin of neural scaling laws: from random graphs to natural language
by: Barkeshli, Maissam, et al.
Published: (2026) -
Generalization through variance: how noise shapes inductive biases in diffusion models
by: Vastola, John J.
Published: (2025)