Stacking as Accelerated Gradient Descent
Fuente:
arXiv
Saved in:
| Main Authors: | Agarwal, Naman, Awasthi, Pranjal, Kale, Satyen, Zhao, Eric |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Agnostic Learning of General ReLU Activation Using Gradient Descent
by: Awasthi, Pranjal, et al.
Published: (2022)
by: Awasthi, Pranjal, et al.
Published: (2022)
Faster Rates For Federated Variational Inequalities
by: Wang, Guanghui, et al.
Published: (2026)
by: Wang, Guanghui, et al.
Published: (2026)
Sample, Scrutinize and Scale: Effective Inference-Time Search by Scaling Verification
by: Zhao, Eric, et al.
Published: (2025)
by: Zhao, Eric, et al.
Published: (2025)
From Style to Facts: Mapping the Boundaries of Knowledge Injection with Finetuning
by: Zhao, Eric, et al.
Published: (2025)
by: Zhao, Eric, et al.
Published: (2025)
The Limits of Preference Data for Post-Training
by: Zhao, Eric, et al.
Published: (2025)
by: Zhao, Eric, et al.
Published: (2025)
AdaBoN: Adaptive Best-of-N Alignment
by: Raman, Vinod, et al.
Published: (2025)
by: Raman, Vinod, et al.
Published: (2025)
Understanding Outer Optimizers in Local SGD: Learning Rates, Momentum, and Acceleration
by: Khaled, Ahmed, et al.
Published: (2025)
by: Khaled, Ahmed, et al.
Published: (2025)
Optimizing Federated Learning using Remote Embeddings for Graph Neural Networks
by: Naman, Pranjal, et al.
Published: (2025)
by: Naman, Pranjal, et al.
Published: (2025)
Exponential Convergence of (Stochastic) Gradient Descent for Separable Logistic Regression
by: Kale, Sacchit, et al.
Published: (2026)
by: Kale, Sacchit, et al.
Published: (2026)
Gradient Dynamics of Attention: How Cross-Entropy Sculpts Bayesian Manifolds
by: Agarwal, Naman, et al.
Published: (2025)
by: Agarwal, Naman, et al.
Published: (2025)
Anytime Acceleration of Gradient Descent
by: Zhang, Zihan, et al.
Published: (2024)
by: Zhang, Zihan, et al.
Published: (2024)
Majority Kernels: An Approach to Leverage Big Model Dynamics for Efficient Small Model Training
by: Mazzawi, Hanna, et al.
Published: (2024)
by: Mazzawi, Hanna, et al.
Published: (2024)
Learning Neural Networks with Sparse Activations
by: Awasthi, Pranjal, et al.
Published: (2024)
by: Awasthi, Pranjal, et al.
Published: (2024)
Accelerated Gradient Descent for Faster Convergence with Minimal Overhead
by: Graca, Manuel, et al.
Published: (2026)
by: Graca, Manuel, et al.
Published: (2026)
Preconditioning for Accelerated Gradient Descent Optimization and Regularization
by: Ye, Qiang
Published: (2024)
by: Ye, Qiang
Published: (2024)
Streaming Krylov-Accelerated Stochastic Gradient Descent
by: Thomas, Stephen
Published: (2025)
by: Thomas, Stephen
Published: (2025)
Product-Stability: Provable Convergence for Gradient Descent on the Edge of Stability
by: Gan, Eric
Published: (2026)
by: Gan, Eric
Published: (2026)
Large Stepsizes Accelerate Gradient Descent for Regularized Logistic Regression
by: Wu, Jingfeng, et al.
Published: (2025)
by: Wu, Jingfeng, et al.
Published: (2025)
Egalitarian Gradient Descent: A Simple Approach to Accelerated Grokking
by: Pasand, Ali Saheb, et al.
Published: (2025)
by: Pasand, Ali Saheb, et al.
Published: (2025)
Can Gradient Descent Simulate Prompting?
by: Zhang, Eric, et al.
Published: (2025)
by: Zhang, Eric, et al.
Published: (2025)
Sample-Efficient Optimization over Generative Priors via Coarse Learnability
by: Awasthi, Pranjal, et al.
Published: (2025)
by: Awasthi, Pranjal, et al.
Published: (2025)
Accelerating Convergence of Stein Variational Gradient Descent via Deep Unfolding
by: Kawamura, Yuya, et al.
Published: (2024)
by: Kawamura, Yuya, et al.
Published: (2024)
Faster Convergence of Stochastic Accelerated Gradient Descent under Interpolation
by: Mishkin, Aaron, et al.
Published: (2024)
by: Mishkin, Aaron, et al.
Published: (2024)
Occam Gradient Descent
by: Kausik, B. N.
Published: (2024)
by: Kausik, B. N.
Published: (2024)
DemOpts: Fairness corrections in COVID-19 case prediction models
by: Awasthi, Naman, et al.
Published: (2024)
by: Awasthi, Naman, et al.
Published: (2024)
Accelerating Feedback-based Algorithms for Quantum Optimization Using Gradient Descent
by: Mozakka, Masih, et al.
Published: (2026)
by: Mozakka, Masih, et al.
Published: (2026)
Towards Noise-adaptive, Problem-adaptive (Accelerated) Stochastic Gradient Descent
by: Vaswani, Sharan, et al.
Published: (2021)
by: Vaswani, Sharan, et al.
Published: (2021)
Accelerating Natural Gradient Descent for PINNs with Randomized Numerical Linear Algebra
by: Bioli, Ivan, et al.
Published: (2025)
by: Bioli, Ivan, et al.
Published: (2025)
First and Second Order Approximations to Stochastic Gradient Descent Methods with Momentum Terms
by: Lu, Eric
Published: (2025)
by: Lu, Eric
Published: (2025)
Efficient Stagewise Pretraining via Progressive Subnetworks
by: Panigrahi, Abhishek, et al.
Published: (2024)
by: Panigrahi, Abhishek, et al.
Published: (2024)
Enhancing Stochastic Gradient Descent: A Unified Framework and Novel Acceleration Methods for Faster Convergence
by: Deng, Yichuan, et al.
Published: (2024)
by: Deng, Yichuan, et al.
Published: (2024)
Increasing Both Batch Size and Learning Rate Accelerates Stochastic Gradient Descent
by: Umeda, Hikaru, et al.
Published: (2024)
by: Umeda, Hikaru, et al.
Published: (2024)
Stochastic Adaptive Gradient Descent Without Descent
by: Aujol, Jean-François, et al.
Published: (2025)
by: Aujol, Jean-François, et al.
Published: (2025)
Corner Gradient Descent
by: Yarotsky, Dmitry
Published: (2025)
by: Yarotsky, Dmitry
Published: (2025)
Gradient Descent as a Perceptron Algorithm: Understanding Dynamics and Implicit Acceleration
by: Tyurin, Alexander
Published: (2025)
by: Tyurin, Alexander
Published: (2025)
Auditing the Fairness of the US COVID-19 Forecast Hub's Case Prediction Models
by: Abrar, Saad Mohammad, et al.
Published: (2024)
by: Abrar, Saad Mohammad, et al.
Published: (2024)
Robust Gradient Descent for Phase Retrieval
by: Buna, Alex, et al.
Published: (2024)
by: Buna, Alex, et al.
Published: (2024)
Distributed Gradient Descent for Functional Learning
by: Yu, Zhan, et al.
Published: (2023)
by: Yu, Zhan, et al.
Published: (2023)
On the Generalization of Stochastic Gradient Descent with Momentum
by: Ramezani-Kebrya, Ali, et al.
Published: (2018)
by: Ramezani-Kebrya, Ali, et al.
Published: (2018)
FedBCD:Communication-Efficient Accelerated Block Coordinate Gradient Descent for Federated Learning
by: Liu, Junkang, et al.
Published: (2026)
by: Liu, Junkang, et al.
Published: (2026)
Similar Items
-
Agnostic Learning of General ReLU Activation Using Gradient Descent
by: Awasthi, Pranjal, et al.
Published: (2022) -
Faster Rates For Federated Variational Inequalities
by: Wang, Guanghui, et al.
Published: (2026) -
Sample, Scrutinize and Scale: Effective Inference-Time Search by Scaling Verification
by: Zhao, Eric, et al.
Published: (2025) -
From Style to Facts: Mapping the Boundaries of Knowledge Injection with Finetuning
by: Zhao, Eric, et al.
Published: (2025) -
The Limits of Preference Data for Post-Training
by: Zhao, Eric, et al.
Published: (2025)