Memory capacity of two layer neural networks with smooth activations
Fuente:
arXiv
Guardado en:
| Autores principales: | Madden, Liam, Thrampoulidis, Christos |
|---|---|
| Formato: | Preprint |
| Publicado: |
2023
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Interpolation with deep neural networks with non-polynomial activations: necessary and sufficient numbers of neurons
por: Madden, Liam
Publicado: (2024)
por: Madden, Liam
Publicado: (2024)
Next-token prediction capacity: general upper bounds and a lower bound for transformers
por: Madden, Liam, et al.
Publicado: (2024)
por: Madden, Liam, et al.
Publicado: (2024)
Escaping Local Minima Provably in Non-convex Matrix Sensing: A Deterministic Framework via Simulated Lifting
por: Shen, Tianqi, et al.
Publicado: (2026)
por: Shen, Tianqi, et al.
Publicado: (2026)
Koopman Lifted Finite Memory Identification via Truncated Grunwald Letnikov Kernels
por: Mojahed, Navid, et al.
Publicado: (2026)
por: Mojahed, Navid, et al.
Publicado: (2026)
Smooth Integer Encoding via Integral Balance
por: Semenov, Stanislav
Publicado: (2025)
por: Semenov, Stanislav
Publicado: (2025)
A smoothing proximal gradient algorithm for matrix rank minimization problem
por: Yu, Quan, et al.
Publicado: (2021)
por: Yu, Quan, et al.
Publicado: (2021)
Using orthogonally structured positive bases for constructing positive $k$-spanning sets with cosine measure guarantees
por: Hare, Warren, et al.
Publicado: (2023)
por: Hare, Warren, et al.
Publicado: (2023)
On Bias and Its Reduction via Standardization in Discretized Electromagnetic Source Localization Problems
por: Lahtinen, Joonas
Publicado: (2024)
por: Lahtinen, Joonas
Publicado: (2024)
Comments and extensions on "State-equivalent form and minimum-order compensator design for rectangular descriptor systems"
por: Shi, Shuo, et al.
Publicado: (2025)
por: Shi, Shuo, et al.
Publicado: (2025)
A Complete Loss Landscape Analysis of Regularized Deep Matrix Factorization
por: Chen, Po, et al.
Publicado: (2025)
por: Chen, Po, et al.
Publicado: (2025)
Learning time-scales in two-layers neural networks
por: Berthier, Raphaël, et al.
Publicado: (2023)
por: Berthier, Raphaël, et al.
Publicado: (2023)
Generic linear convergence for algorithms of non-linear least squares over smooth varieties
por: Hu, Shenglong, et al.
Publicado: (2025)
por: Hu, Shenglong, et al.
Publicado: (2025)
Implicit Bias of Spectral Descent and Muon on Multiclass Separable Data
por: Fan, Chen, et al.
Publicado: (2025)
por: Fan, Chen, et al.
Publicado: (2025)
Implicit Bias and Fast Convergence Rates for Self-attention
por: Vasudeva, Bhavya, et al.
Publicado: (2024)
por: Vasudeva, Bhavya, et al.
Publicado: (2024)
SAD Neural Networks: Divergent Gradient Flows and Asymptotic Optimality via o-minimal Structures
por: Kranz, Julian, et al.
Publicado: (2025)
por: Kranz, Julian, et al.
Publicado: (2025)
Efficient Defection: Overage-Proportional Rationing Attains the Cooperative Frontier
por: Lengyel, Florian
Publicado: (2025)
por: Lengyel, Florian
Publicado: (2025)
Optimization on the Oblique Manifold for Sparse Simplex Constraints via Multiplicative Updates
por: Esposito, Flavia, et al.
Publicado: (2025)
por: Esposito, Flavia, et al.
Publicado: (2025)
Diagonalizing the Softmax: Hadamard Initialization for Tractable Cross-Entropy Dynamics
por: Garrod, Connall, et al.
Publicado: (2025)
por: Garrod, Connall, et al.
Publicado: (2025)
High-precision linear minimization is no slower than projection
por: Woodstock, Zev
Publicado: (2025)
por: Woodstock, Zev
Publicado: (2025)
Hautus-Type Criteria for Controllability and Stabilizability of Backward-Structured Stochastic Systems
por: Sun, Jingrui
Publicado: (2026)
por: Sun, Jingrui
Publicado: (2026)
Local minima in Newton's aerodynamical problem and inequalities between norms of partial derivatives
por: Plakhov, Alexander, et al.
Publicado: (2024)
por: Plakhov, Alexander, et al.
Publicado: (2024)
Boundary Regional Controllability of Semilinear Systems Involving Caputo Time Fractional Derivatives
por: Tajani, Asmae, et al.
Publicado: (2024)
por: Tajani, Asmae, et al.
Publicado: (2024)
A Normal Map-Based Proximal Stochastic Gradient Method: Convergence and Identification Properties
por: Qiu, Junwen, et al.
Publicado: (2023)
por: Qiu, Junwen, et al.
Publicado: (2023)
A New Random Reshuffling Method for Nonsmooth Nonconvex Finite-sum Optimization
por: Qiu, Junwen, et al.
Publicado: (2023)
por: Qiu, Junwen, et al.
Publicado: (2023)
Optimization without Retraction on the Random Generalized Stiefel Manifold
por: Vary, Simon, et al.
Publicado: (2024)
por: Vary, Simon, et al.
Publicado: (2024)
Convergence of gradient descent for deep neural networks
por: Chatterjee, Sourav
Publicado: (2022)
por: Chatterjee, Sourav
Publicado: (2022)
On semimonotone matrices of exact order two
por: Chauhan, Bharat Pratap, et al.
Publicado: (2026)
por: Chauhan, Bharat Pratap, et al.
Publicado: (2026)
On the Characterization of gH-partial derivatives and gH-Product for Interval-Valued Functions
por: Suhail, Amir, et al.
Publicado: (2025)
por: Suhail, Amir, et al.
Publicado: (2025)
On the Optimization and Generalization of Multi-head Attention
por: Deora, Puneesh, et al.
Publicado: (2023)
por: Deora, Puneesh, et al.
Publicado: (2023)
Consensus-based optimization for closed-box adversarial attacks and a connection to evolution strategies
por: Roith, Tim, et al.
Publicado: (2025)
por: Roith, Tim, et al.
Publicado: (2025)
The Fundamental Theorem of Calculus in higher dimensions
por: Bár, Filip
Publicado: (2024)
por: Bár, Filip
Publicado: (2024)
Stiefel optimization is NP-hard
por: Lai, Zehua, et al.
Publicado: (2025)
por: Lai, Zehua, et al.
Publicado: (2025)
A Large Deviations Perspective on Policy Gradient Algorithms
por: Jongeneel, Wouter, et al.
Publicado: (2023)
por: Jongeneel, Wouter, et al.
Publicado: (2023)
Ultra-fast feature learning for the training of two-layer neural networks in the two-timescale regime
por: Barboni, Raphaël, et al.
Publicado: (2025)
por: Barboni, Raphaël, et al.
Publicado: (2025)
Roots, trace, and extendability of flat nonnegative smooth functions
por: Jiang, Fushuai
Publicado: (2023)
por: Jiang, Fushuai
Publicado: (2023)
On the controllability of nonlinear systems with a periodic drift
por: Caillau, Jean-Baptiste, et al.
Publicado: (2024)
por: Caillau, Jean-Baptiste, et al.
Publicado: (2024)
Instability and Efficiency of Non-cooperative Games
por: Zhang, Jianfeng
Publicado: (2024)
por: Zhang, Jianfeng
Publicado: (2024)
Ultracoarse Equilibria and Ordinal-Folding Dynamics in Operator-Algebraic Models of Infinite Multi-Agent Games
por: Alpay, Faruk, et al.
Publicado: (2025)
por: Alpay, Faruk, et al.
Publicado: (2025)
FedSLoP: Memory-Efficient Federated Learning with Low-Rank Gradient Projection
por: He, Yutong, et al.
Publicado: (2026)
por: He, Yutong, et al.
Publicado: (2026)
Monotonicity of the jump set and jump amplitudes in one-dimensional TV denoising
por: Cristoferi, Riccardo, et al.
Publicado: (2025)
por: Cristoferi, Riccardo, et al.
Publicado: (2025)
Ejemplares similares
-
Interpolation with deep neural networks with non-polynomial activations: necessary and sufficient numbers of neurons
por: Madden, Liam
Publicado: (2024) -
Next-token prediction capacity: general upper bounds and a lower bound for transformers
por: Madden, Liam, et al.
Publicado: (2024) -
Escaping Local Minima Provably in Non-convex Matrix Sensing: A Deterministic Framework via Simulated Lifting
por: Shen, Tianqi, et al.
Publicado: (2026) -
Koopman Lifted Finite Memory Identification via Truncated Grunwald Letnikov Kernels
por: Mojahed, Navid, et al.
Publicado: (2026) -
Smooth Integer Encoding via Integral Balance
por: Semenov, Stanislav
Publicado: (2025)