Saved in:
| Main Authors: | Grishina, Ekaterina, Gorbunov, Mikhail, Rakhuba, Maxim |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2409.11859 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ProcrustesGPT: Compressing LLMs with Structured Matrices and Orthogonal Transformations
by: Grishina, Ekaterina, et al.
Published: (2025)
by: Grishina, Ekaterina, et al.
Published: (2025)
On the Upper Bounds for the Matrix Spectral Norm
by: Naumov, Alexey, et al.
Published: (2025)
by: Naumov, Alexey, et al.
Published: (2025)
DyKAF: Dynamical Kronecker Approximation of the Fisher Information Matrix for Gradient Preconditioning
by: Yudin, Nikolay, et al.
Published: (2025)
by: Yudin, Nikolay, et al.
Published: (2025)
Group and Shuffle: Efficient Structured Orthogonal Parametrization
by: Gorbunov, Mikhail, et al.
Published: (2024)
by: Gorbunov, Mikhail, et al.
Published: (2024)
Accelerating Newton-Schulz Iteration for Orthogonalization via Chebyshev-type Polynomials
by: Grishina, Ekaterina, et al.
Published: (2025)
by: Grishina, Ekaterina, et al.
Published: (2025)
Matrix-Free Two-to-Infinity and One-to-Two Norms Estimation
by: Tsyganov, Askar, et al.
Published: (2025)
by: Tsyganov, Askar, et al.
Published: (2025)
COALA: Numerically Stable and Efficient Framework for Context-Aware Low-Rank Approximation
by: Parkina, Uliana, et al.
Published: (2025)
by: Parkina, Uliana, et al.
Published: (2025)
Spectral Norm of Convolutional Layers with Circular and Zero Paddings
by: Delattre, Blaise, et al.
Published: (2024)
by: Delattre, Blaise, et al.
Published: (2024)
Ultra Fast Warm Start Solution for Graph Recommendations
by: Yusupov, Viacheslav, et al.
Published: (2025)
by: Yusupov, Viacheslav, et al.
Published: (2025)
Leveraging Geometric Insights in Hyperbolic Triplet Loss for Improved Recommendations
by: Yusupov, Viacheslav, et al.
Published: (2025)
by: Yusupov, Viacheslav, et al.
Published: (2025)
Pay Attention to Attention Distribution: A New Local Lipschitz Bound for Transformers
by: Yudin, Nikolay, et al.
Published: (2025)
by: Yudin, Nikolay, et al.
Published: (2025)
Knowledge Graph Completion with Mixed Geometry Tensor Factorization
by: Yusupov, Viacheslav, et al.
Published: (2025)
by: Yusupov, Viacheslav, et al.
Published: (2025)
Reliable Abstention under Adversarial Injections: Tight Lower Bounds and New Upper Bounds
by: Edelman, Ezra, et al.
Published: (2026)
by: Edelman, Ezra, et al.
Published: (2026)
On Stopping Times of Power-one Sequential Tests: Tight Lower and Upper Bounds
by: Agrawal, Shubhada, et al.
Published: (2025)
by: Agrawal, Shubhada, et al.
Published: (2025)
Clipping Improves Adam-Norm and AdaGrad-Norm when the Noise Is Heavy-Tailed
by: Chezhegov, Savelii, et al.
Published: (2024)
by: Chezhegov, Savelii, et al.
Published: (2024)
Trade-off Functions for DP-SGD with Subsampling based on Random Shuffling: Tight Upper and Lower Bounds
by: van Dijk, Marten, et al.
Published: (2026)
by: van Dijk, Marten, et al.
Published: (2026)
An Optimal Tightness Bound for the Simulation Lemma
by: Lobel, Sam, et al.
Published: (2024)
by: Lobel, Sam, et al.
Published: (2024)
Tight Regret Upper and Lower Bounds for Optimistic Hedge in Two-Player Zero-Sum Games
by: Tsuchiya, Taira
Published: (2025)
by: Tsuchiya, Taira
Published: (2025)
Geometry and Dynamics of LayerNorm
by: Riechers, Paul M.
Published: (2024)
by: Riechers, Paul M.
Published: (2024)
Tight Convergence Rate Bounds for Optimization Under Power Law Spectral Conditions
by: Velikanov, Maksim, et al.
Published: (2022)
by: Velikanov, Maksim, et al.
Published: (2022)
Norm-Bounded Low-Rank Adaptation
by: Wang, Ruigang, et al.
Published: (2025)
by: Wang, Ruigang, et al.
Published: (2025)
Which Algorithms Have Tight Generalization Bounds?
by: Gastpar, Michael, et al.
Published: (2024)
by: Gastpar, Michael, et al.
Published: (2024)
SDP-CROWN: Efficient Bound Propagation for Neural Network Verification with Tightness of Semidefinite Programming
by: Chiu, Hong-Ming, et al.
Published: (2025)
by: Chiu, Hong-Ming, et al.
Published: (2025)
Upper Bounds for Local Learning Coefficients of Three-Layer Neural Networks
by: Kurumadani, Yuki
Published: (2026)
by: Kurumadani, Yuki
Published: (2026)
Tight Non-asymptotic Inference via Sub-Gaussian Intrinsic Moment Norm
by: Zhang, Huiming, et al.
Published: (2023)
by: Zhang, Huiming, et al.
Published: (2023)
E-Globe: Scalable $ε$-Global Verification of Neural Networks via Tight Upper Bounds and Pattern-Aware Branching
by: Li, Wenting, et al.
Published: (2026)
by: Li, Wenting, et al.
Published: (2026)
Concept Drift Guided LayerNorm Tuning for Efficient Multimodal Metaphor Identification
by: Qian, Wenhao, et al.
Published: (2025)
by: Qian, Wenhao, et al.
Published: (2025)
Tight Bounds for Jensen's Gap with Applications to Variational Inference
by: Mazur, Marcin, et al.
Published: (2025)
by: Mazur, Marcin, et al.
Published: (2025)
Layer-wise Adaptive Gradient Norm Penalizing Method for Efficient and Accurate Deep Learning
by: Lee, Sunwoo
Published: (2025)
by: Lee, Sunwoo
Published: (2025)
On the Role of Attention Masks and LayerNorm in Transformers
by: Wu, Xinyi, et al.
Published: (2024)
by: Wu, Xinyi, et al.
Published: (2024)
When Does Removing LayerNorm Help? Activation Bounding as a Regime-Dependent Implicit Regularizer
by: Verma, Lucky
Published: (2026)
by: Verma, Lucky
Published: (2026)
Tight Generalization Bounds for Noiseless Inverse Optimization
by: Fatemi, Pouria, et al.
Published: (2026)
by: Fatemi, Pouria, et al.
Published: (2026)
Tight Generalization Bounds for Large-Margin Halfspaces
by: Larsen, Kasper Green, et al.
Published: (2025)
by: Larsen, Kasper Green, et al.
Published: (2025)
High Probability Complexity Bounds for Non-Smooth Stochastic Optimization with Heavy-Tailed Noise
by: Gorbunov, Eduard, et al.
Published: (2021)
by: Gorbunov, Eduard, et al.
Published: (2021)
Tight Sample Complexity Bounds for Entropic Best Policy Identification
by: Essakine, Amer, et al.
Published: (2026)
by: Essakine, Amer, et al.
Published: (2026)
Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability
by: Baroni, Luca, et al.
Published: (2025)
by: Baroni, Luca, et al.
Published: (2025)
Tight Bounds for Learning Polyhedra with a Margin
by: Patel, Shyamal, et al.
Published: (2026)
by: Patel, Shyamal, et al.
Published: (2026)
Muon Optimizes Under Spectral Norm Constraints
by: Chen, Lizhang, et al.
Published: (2025)
by: Chen, Lizhang, et al.
Published: (2025)
Just One Layer Norm Guarantees Stable Extrapolation
by: Ziomek, Juliusz, et al.
Published: (2025)
by: Ziomek, Juliusz, et al.
Published: (2025)
Tight Lower Bounds and Improved Convergence in Performative Prediction
by: Khorsandi, Pedram, et al.
Published: (2024)
by: Khorsandi, Pedram, et al.
Published: (2024)
Similar Items
-
ProcrustesGPT: Compressing LLMs with Structured Matrices and Orthogonal Transformations
by: Grishina, Ekaterina, et al.
Published: (2025) -
On the Upper Bounds for the Matrix Spectral Norm
by: Naumov, Alexey, et al.
Published: (2025) -
DyKAF: Dynamical Kronecker Approximation of the Fisher Information Matrix for Gradient Preconditioning
by: Yudin, Nikolay, et al.
Published: (2025) -
Group and Shuffle: Efficient Structured Orthogonal Parametrization
by: Gorbunov, Mikhail, et al.
Published: (2024) -
Accelerating Newton-Schulz Iteration for Orthogonalization via Chebyshev-type Polynomials
by: Grishina, Ekaterina, et al.
Published: (2025)