Implicit Bias in Matrix Factorization and its Explicit Realization in a New Architecture
Fuente:
arXiv
Saved in:
| Main Authors: | Hou, Yikun, Sra, Suvrit, Yurtsever, Alp |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Revisiting Frank-Wolfe for Structured Nonconvex Optimization
by: Maskan, Hoomaan, et al.
Published: (2025)
by: Maskan, Hoomaan, et al.
Published: (2025)
Linearly Convergent Algorithms for Nonsmooth Problems with Unknown Smooth Pieces
by: Zhang, Zhe, et al.
Published: (2025)
by: Zhang, Zhe, et al.
Published: (2025)
Toward generalizable learning of all (linear) first-order methods via memory augmented Transformers
by: Dutta, Sanchayan, et al.
Published: (2024)
by: Dutta, Sanchayan, et al.
Published: (2024)
The Multi-Block DC Function Class: Theory, Algorithms, and Applications
by: Fatemi, Pouria, et al.
Published: (2026)
by: Fatemi, Pouria, et al.
Published: (2026)
Improved Rates for Stochastic Variance-Reduced Difference-of-Convex Algorithms
by: Nguyen, Anh Duc, et al.
Published: (2025)
by: Nguyen, Anh Duc, et al.
Published: (2025)
Randomized Block Coordinate DC Programming
by: Maskan, Hoomaan, et al.
Published: (2024)
by: Maskan, Hoomaan, et al.
Published: (2024)
How to escape sharp minima with random perturbations
by: Ahn, Kwangjun, et al.
Published: (2023)
by: Ahn, Kwangjun, et al.
Published: (2023)
Riemannian Bilevel Optimization
by: Dutta, Sanchayan, et al.
Published: (2024)
by: Dutta, Sanchayan, et al.
Published: (2024)
Cost-Driven Representation Learning for Linear Quadratic Gaussian Control: Part I
by: Tian, Yi, et al.
Published: (2022)
by: Tian, Yi, et al.
Published: (2022)
Cost-Driven Representation Learning for Linear Quadratic Gaussian Control: Part II
by: Tian, Yi, et al.
Published: (2026)
by: Tian, Yi, et al.
Published: (2026)
Convex Formulations for Training Two-Layer ReLU Neural Networks
by: Prakhya, Karthik, et al.
Published: (2024)
by: Prakhya, Karthik, et al.
Published: (2024)
Tight Generalization Bounds for Noiseless Inverse Optimization
by: Fatemi, Pouria, et al.
Published: (2026)
by: Fatemi, Pouria, et al.
Published: (2026)
First-Order Methods for Linearly Constrained Bilevel Optimization
by: Kornowski, Guy, et al.
Published: (2024)
by: Kornowski, Guy, et al.
Published: (2024)
Provable Reduction in Communication Rounds for Non-Smooth Convex Federated Learning
by: Palenzuela, Karlo, et al.
Published: (2025)
by: Palenzuela, Karlo, et al.
Published: (2025)
Implicit Bias and Convergence of Matrix Stochastic Mirror Descent
by: Akhtiamov, Danil, et al.
Published: (2026)
by: Akhtiamov, Danil, et al.
Published: (2026)
Personalized Multi-tier Federated Learning
by: Banerjee, Sourasekhar, et al.
Published: (2024)
by: Banerjee, Sourasekhar, et al.
Published: (2024)
Linear attention is (maybe) all you need (to understand transformer optimization)
by: Ahn, Kwangjun, et al.
Published: (2023)
by: Ahn, Kwangjun, et al.
Published: (2023)
Exact Convex Reformulations of Linear Neural Networks via Completely Positive Lifting
by: Prakhya, Karthik, et al.
Published: (2026)
by: Prakhya, Karthik, et al.
Published: (2026)
The Implicit Bias of Heterogeneity towards Invariance: A Study of Multi-Environment Matrix Sensing
by: Xu, Yang, et al.
Published: (2024)
by: Xu, Yang, et al.
Published: (2024)
Implicit Regularization in Perturbed Deep Matrix Factorization: Spectral Conditions and Stability
by: Wang, Jingzhe, et al.
Published: (2026)
by: Wang, Jingzhe, et al.
Published: (2026)
Universal Adaptive Proximal Gradient Methods via Gradient Mapping Accumulation
by: Wang, Zimeng, et al.
Published: (2026)
by: Wang, Zimeng, et al.
Published: (2026)
Generalized Stochastic Gradient Descent with Momentum Methods for Smooth Optimization
by: Wang, Zimeng, et al.
Published: (2026)
by: Wang, Zimeng, et al.
Published: (2026)
The Rich and the Simple: On the Implicit Bias of Adam and SGD
by: Vasudeva, Bhavya, et al.
Published: (2025)
by: Vasudeva, Bhavya, et al.
Published: (2025)
Implicit Bias of Mirror Flow on Separable Data
by: Pesme, Scott, et al.
Published: (2024)
by: Pesme, Scott, et al.
Published: (2024)
Implicit Bias and Fast Convergence Rates for Self-attention
by: Vasudeva, Bhavya, et al.
Published: (2024)
by: Vasudeva, Bhavya, et al.
Published: (2024)
Sinkhorn algorithms and linear programming solvers for optimal partial transport problems
by: Bai, Yikun
Published: (2024)
by: Bai, Yikun
Published: (2024)
Implicit Bias of Gradient Descent for Non-Homogeneous Deep Networks
by: Cai, Yuhang, et al.
Published: (2025)
by: Cai, Yuhang, et al.
Published: (2025)
Implicit Bias of Spectral Descent and Muon on Multiclass Separable Data
by: Fan, Chen, et al.
Published: (2025)
by: Fan, Chen, et al.
Published: (2025)
Convergence and Implicit Bias of Gradient Descent on Continual Linear Classification
by: Jung, Hyunji, et al.
Published: (2025)
by: Jung, Hyunji, et al.
Published: (2025)
Towards The Implicit Bias on Multiclass Separable Data Under Norm Constraints
by: Xie, Shengping, et al.
Published: (2026)
by: Xie, Shengping, et al.
Published: (2026)
Implicit Bias of AdamW: $\ell_\infty$ Norm Constrained Optimization
by: Xie, Shuo, et al.
Published: (2024)
by: Xie, Shuo, et al.
Published: (2024)
Implicit Regularization Makes Overparameterized Asymmetric Matrix Sensing Robust to Perturbations
by: Wind, Johan S.
Published: (2023)
by: Wind, Johan S.
Published: (2023)
On the Implicit Bias of Adam
by: Cattaneo, Matias D., et al.
Published: (2023)
by: Cattaneo, Matias D., et al.
Published: (2023)
Sharpness of Minima in Deep Matrix Factorization
by: Kamber, Anil, et al.
Published: (2025)
by: Kamber, Anil, et al.
Published: (2025)
Sum-of-norms regularized Nonnegative Matrix Factorization
by: Ang, Andersen, et al.
Published: (2024)
by: Ang, Andersen, et al.
Published: (2024)
Convergence of Alternating Gradient Descent for Matrix Factorization
by: Ward, Rachel, et al.
Published: (2023)
by: Ward, Rachel, et al.
Published: (2023)
Fast and Effective Computation of Generalized Symmetric Matrix Factorization
by: Yang, Lei, et al.
Published: (2026)
by: Yang, Lei, et al.
Published: (2026)
How Does the ReLU Activation Affect the Implicit Bias of Gradient Descent on High-dimensional Neural Network Regression?
by: Lai, Kuo-Wei, et al.
Published: (2026)
by: Lai, Kuo-Wei, et al.
Published: (2026)
Preconditioned Gradient Descent for Over-Parameterized Nonconvex Matrix Factorization
by: Zhang, Gavin, et al.
Published: (2025)
by: Zhang, Gavin, et al.
Published: (2025)
A Second-Order Majorant Algorithm for Nonnegative Matrix Factorization
by: Pham, Mai-Quyen, et al.
Published: (2023)
by: Pham, Mai-Quyen, et al.
Published: (2023)
Similar Items
-
Revisiting Frank-Wolfe for Structured Nonconvex Optimization
by: Maskan, Hoomaan, et al.
Published: (2025) -
Linearly Convergent Algorithms for Nonsmooth Problems with Unknown Smooth Pieces
by: Zhang, Zhe, et al.
Published: (2025) -
Toward generalizable learning of all (linear) first-order methods via memory augmented Transformers
by: Dutta, Sanchayan, et al.
Published: (2024) -
The Multi-Block DC Function Class: Theory, Algorithms, and Applications
by: Fatemi, Pouria, et al.
Published: (2026) -
Improved Rates for Stochastic Variance-Reduced Difference-of-Convex Algorithms
by: Nguyen, Anh Duc, et al.
Published: (2025)