Saved in:
| Main Author: | Farag, Peter |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2512.23905 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Loss Landscape Degeneracy and Stagewise Development in Transformers
by: Hoogland, Jesse, et al.
Published: (2024)
by: Hoogland, Jesse, et al.
Published: (2024)
Benign Overfitting without Linearity: Neural Network Classifiers Trained by Gradient Descent for Noisy Linear Data
by: Frei, Spencer, et al.
Published: (2022)
by: Frei, Spencer, et al.
Published: (2022)
Transformer-based Stagewise Decomposition for Large-Scale Multistage Stochastic Optimization
by: Kim, Chanyeong, et al.
Published: (2024)
by: Kim, Chanyeong, et al.
Published: (2024)
Stagewise Boosting Distributional Regression
by: Wetscher, Mattias, et al.
Published: (2024)
by: Wetscher, Mattias, et al.
Published: (2024)
Influence Dynamics and Stagewise Data Attribution
by: Lee, Jin Hwa, et al.
Published: (2025)
by: Lee, Jin Hwa, et al.
Published: (2025)
Efficient VQ-QAT and Mixed Vector/Linear quantized Neural Networks
by: Gou, Terry, et al.
Published: (2026)
by: Gou, Terry, et al.
Published: (2026)
SENTINEL: Stagewise Integrity Verification for Pipeline Parallel Decentralized Training
by: Dolatabadi, Hadi Mohaghegh, et al.
Published: (2026)
by: Dolatabadi, Hadi Mohaghegh, et al.
Published: (2026)
Stagewise Reinforcement Learning and the Geometry of the Regret Landscape
by: Elliott, Chris, et al.
Published: (2026)
by: Elliott, Chris, et al.
Published: (2026)
A Generalization Bound for Nearly-Linear Networks
by: Golikov, Eugene
Published: (2024)
by: Golikov, Eugene
Published: (2024)
Learning Linear Utility Functions From Pairwise Comparison Queries
by: Ge, Luise, et al.
Published: (2024)
by: Ge, Luise, et al.
Published: (2024)
Efficient Stagewise Pretraining via Progressive Subnetworks
by: Panigrahi, Abhishek, et al.
Published: (2024)
by: Panigrahi, Abhishek, et al.
Published: (2024)
Expectation Error Bounds for Transfer Learning in Linear Regression and Linear Neural Networks
by: Liu, Meitong, et al.
Published: (2026)
by: Liu, Meitong, et al.
Published: (2026)
Majorization-Minimization Dual Stagewise Algorithm for Generalized Lasso
by: Chen, Jianmin, et al.
Published: (2025)
by: Chen, Jianmin, et al.
Published: (2025)
Gated Linear Attention Transformers with Hardware-Efficient Training
by: Yang, Songlin, et al.
Published: (2023)
by: Yang, Songlin, et al.
Published: (2023)
Who Said Neural Networks Aren't Linear?
by: Berman, Nimrod, et al.
Published: (2025)
by: Berman, Nimrod, et al.
Published: (2025)
An Uncertainty Principle for Linear Recurrent Neural Networks
by: François, Alexandre, et al.
Published: (2025)
by: François, Alexandre, et al.
Published: (2025)
Weight-Space Linear Recurrent Neural Networks
by: Nzoyem, Roussel Desmond, et al.
Published: (2025)
by: Nzoyem, Roussel Desmond, et al.
Published: (2025)
Training Dynamics of In-Context Learning in Linear Attention
by: Zhang, Yedi, et al.
Published: (2025)
by: Zhang, Yedi, et al.
Published: (2025)
Linearly Constrained Weights: Reducing Activation Shift for Faster Training of Neural Networks
by: Kutsuna, Takuro
Published: (2024)
by: Kutsuna, Takuro
Published: (2024)
Multi-Objective Linear Ensembles for Robust and Sparse Training of Few-Bit Neural Networks
by: Bernardelli, Ambrogio Maria, et al.
Published: (2022)
by: Bernardelli, Ambrogio Maria, et al.
Published: (2022)
Optimal Mixed Integer Linear Optimization Trained Multivariate Classification Trees
by: Alston, Brandon, et al.
Published: (2024)
by: Alston, Brandon, et al.
Published: (2024)
Enhancing Deep Neural Network Training Efficiency and Performance through Linear Prediction
by: Ying, Hejie, et al.
Published: (2023)
by: Ying, Hejie, et al.
Published: (2023)
Mixed Dynamics In Linear Networks: Unifying the Lazy and Active Regimes
by: Tu, Zhenfeng, et al.
Published: (2024)
by: Tu, Zhenfeng, et al.
Published: (2024)
Continuous-Time Piecewise-Linear Recurrent Neural Networks
by: Brändle, Alena, et al.
Published: (2026)
by: Brändle, Alena, et al.
Published: (2026)
A Novel Explanation Against Linear Neural Networks
by: Lakkapragada, Anish
Published: (2023)
by: Lakkapragada, Anish
Published: (2023)
Rethinking Transformer Connectivity: TLinFormer, A Path to Exact, Full Context-Aware Linear Attention
by: Tang, Zhongpan
Published: (2025)
by: Tang, Zhongpan
Published: (2025)
The Geometry of the Set of Equivalent Linear Neural Networks
by: Shewchuk, Jonathan Richard, et al.
Published: (2024)
by: Shewchuk, Jonathan Richard, et al.
Published: (2024)
Linear Mode Connectivity in Sparse Neural Networks
by: McDermott, Luke, et al.
Published: (2023)
by: McDermott, Luke, et al.
Published: (2023)
Nearly Optimal Linear Convergence of Stochastic Primal-Dual Methods for Linear Programming
by: Lu, Haihao, et al.
Published: (2021)
by: Lu, Haihao, et al.
Published: (2021)
BELIEF in Dependence: Leveraging Atomic Linearity in Data Bits for Rethinking Generalized Linear Models
by: Brown, Benjamin, et al.
Published: (2022)
by: Brown, Benjamin, et al.
Published: (2022)
Graph Neural Network-Based Distributed Optimal Control for Linear Networked Systems: An Online Distributed Training Approach
by: Song, Zihao, et al.
Published: (2025)
by: Song, Zihao, et al.
Published: (2025)
Dense Feature Learning via Linear Structure Preservation in Medical Data
by: Zhang, Yuanyun, et al.
Published: (2026)
by: Zhang, Yuanyun, et al.
Published: (2026)
Dense Training, Sparse Inference: Rethinking Training of Mixture-of-Experts Language Models
by: Pan, Bowen, et al.
Published: (2024)
by: Pan, Bowen, et al.
Published: (2024)
Fair Generalized Linear Mixed Models
by: Burgard, Jan Pablo, et al.
Published: (2024)
by: Burgard, Jan Pablo, et al.
Published: (2024)
Combining Graph Neural Networks and Mixed Integer Linear Programming for Molecular Inference under the Two-Layered Model
by: Zhu, Jianshen, et al.
Published: (2025)
by: Zhu, Jianshen, et al.
Published: (2025)
Recurrent Neural Networks with Linear Structures for Electricity Price Forecasting
by: Amor, Souhir Ben, et al.
Published: (2025)
by: Amor, Souhir Ben, et al.
Published: (2025)
Coordinate Encoding on Linear Grids for Physics-Informed Neural Networks
by: Tsuchino, Tetsuro, et al.
Published: (2026)
by: Tsuchino, Tetsuro, et al.
Published: (2026)
Local Linear Recovery Guarantee of Deep Neural Networks at Overparameterization
by: Zhang, Yaoyu, et al.
Published: (2024)
by: Zhang, Yaoyu, et al.
Published: (2024)
Mixture of Linear Models Co-supervised by Deep Neural Networks
by: Seo, Beomseok, et al.
Published: (2021)
by: Seo, Beomseok, et al.
Published: (2021)
Weak-to-Strong Generalization is Nearly Inevitable (in Linear Models)
by: Geng, Scott, et al.
Published: (2026)
by: Geng, Scott, et al.
Published: (2026)
Similar Items
-
Loss Landscape Degeneracy and Stagewise Development in Transformers
by: Hoogland, Jesse, et al.
Published: (2024) -
Benign Overfitting without Linearity: Neural Network Classifiers Trained by Gradient Descent for Noisy Linear Data
by: Frei, Spencer, et al.
Published: (2022) -
Transformer-based Stagewise Decomposition for Large-Scale Multistage Stochastic Optimization
by: Kim, Chanyeong, et al.
Published: (2024) -
Stagewise Boosting Distributional Regression
by: Wetscher, Mattias, et al.
Published: (2024) -
Influence Dynamics and Stagewise Data Attribution
by: Lee, Jin Hwa, et al.
Published: (2025)