SLTrain: a sparse plus low-rank approach for parameter and memory efficient pretraining
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Han, Andi, Li, Jiaxiang, Huang, Wei, Hong, Mingyi, Takeda, Akiko, Jawanpuria, Pratik, Mishra, Bamdev |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Framework for Bilevel Optimization on Riemannian Manifolds
von: Han, Andi, et al.
Veröffentlicht: (2024)
von: Han, Andi, et al.
Veröffentlicht: (2024)
Riemannian coordinate descent algorithms on matrix manifolds
von: Han, Andi, et al.
Veröffentlicht: (2024)
von: Han, Andi, et al.
Veröffentlicht: (2024)
Generalized infinite dimensional Alpha-Procrustes based geometries
von: Goomanee, Salvish, et al.
Veröffentlicht: (2025)
von: Goomanee, Salvish, et al.
Veröffentlicht: (2025)
Riemannian Optimization for Hadamard Products of Low-Rank Matrices
von: Jawanpuria, Pratik, et al.
Veröffentlicht: (2026)
von: Jawanpuria, Pratik, et al.
Veröffentlicht: (2026)
A Gauss-Newton Approach for Min-Max Optimization in Generative Adversarial Networks
von: Mishra, Neel, et al.
Veröffentlicht: (2024)
von: Mishra, Neel, et al.
Veröffentlicht: (2024)
A Riemannian Approach to Ground Metric Learning for Optimal Transport
von: Jawanpuria, Pratik, et al.
Veröffentlicht: (2024)
von: Jawanpuria, Pratik, et al.
Veröffentlicht: (2024)
LOFT: Low-Rank Orthogonal Fine-Tuning via Task-Aware Support Selection
von: Zhao, Lanxin, et al.
Veröffentlicht: (2026)
von: Zhao, Lanxin, et al.
Veröffentlicht: (2026)
Riemannian Federated Learning via Averaging Gradient Streams
von: Huang, Zhenwei, et al.
Veröffentlicht: (2024)
von: Huang, Zhenwei, et al.
Veröffentlicht: (2024)
Federated Learning on Riemannian Manifolds with Differential Privacy
von: Huang, Zhenwei, et al.
Veröffentlicht: (2024)
von: Huang, Zhenwei, et al.
Veröffentlicht: (2024)
Intrinsic Muon: Spectral Optimization on Riemannian Matrix Manifolds
von: Li, Yibang, et al.
Veröffentlicht: (2026)
von: Li, Yibang, et al.
Veröffentlicht: (2026)
Submodular Framework for Structured-Sparse Optimal Transport
von: Manupriya, Piyushi, et al.
Veröffentlicht: (2024)
von: Manupriya, Piyushi, et al.
Veröffentlicht: (2024)
Memory-Efficient LLM Pretraining via Minimalist Optimizer Design
von: Glentis, Athanasios, et al.
Veröffentlicht: (2025)
von: Glentis, Athanasios, et al.
Veröffentlicht: (2025)
UniPROT: Uniform Prototype Selection via Partial Optimal Transport with Submodular Guarantees
von: Chanda, Prateek, et al.
Veröffentlicht: (2026)
von: Chanda, Prateek, et al.
Veröffentlicht: (2026)
Nyström Approximation on Manifolds
von: Nie, Hantao, et al.
Veröffentlicht: (2026)
von: Nie, Hantao, et al.
Veröffentlicht: (2026)
Efficient Optimization with Orthogonality Constraint: a Randomized Riemannian Submanifold Method
von: Han, Andi, et al.
Veröffentlicht: (2025)
von: Han, Andi, et al.
Veröffentlicht: (2025)
Scalable Parameter and Memory Efficient Pretraining for LLM: Recent Algorithmic Advances and Benchmarking
von: Glentis, Athanasios, et al.
Veröffentlicht: (2025)
von: Glentis, Athanasios, et al.
Veröffentlicht: (2025)
Modified K-means Algorithm with Local Optimality Guarantees
von: Li, Mingyi, et al.
Veröffentlicht: (2025)
von: Li, Mingyi, et al.
Veröffentlicht: (2025)
FedSEA: Achieving Benefit of Parallelization in Federated Online Learning
von: Sahu, Harekrushna, et al.
Veröffentlicht: (2026)
von: Sahu, Harekrushna, et al.
Veröffentlicht: (2026)
MMD-Regularized Unbalanced Optimal Transport
von: Manupriya, Piyushi, et al.
Veröffentlicht: (2020)
von: Manupriya, Piyushi, et al.
Veröffentlicht: (2020)
On the Role of Label Noise in the Feature Learning Process
von: Han, Andi, et al.
Veröffentlicht: (2025)
von: Han, Andi, et al.
Veröffentlicht: (2025)
AdaFish: Fast low-rank parameter-efficient fine-tuning by using second-order information
von: Hu, Jiang, et al.
Veröffentlicht: (2024)
von: Hu, Jiang, et al.
Veröffentlicht: (2024)
A Framework for Quantifying How Pre-Training and Context Benefit In-Context Learning
von: Song, Bingqing, et al.
Veröffentlicht: (2025)
von: Song, Bingqing, et al.
Veröffentlicht: (2025)
Beyond IID weights: sparse and low-rank deep Neural Networks are also Gaussian Processes
von: Nait-Saada, Thiziri, et al.
Veröffentlicht: (2023)
von: Nait-Saada, Thiziri, et al.
Veröffentlicht: (2023)
Do pretrained Transformers Learn In-Context by Gradient Descent?
von: Shen, Lingfeng, et al.
Veröffentlicht: (2023)
von: Shen, Lingfeng, et al.
Veröffentlicht: (2023)
Adaptive regularization parameter selection for high-dimensional inverse problems: A Bayesian approach with Tucker low-rank constraints
von: Yang, Qing-Mei, et al.
Veröffentlicht: (2026)
von: Yang, Qing-Mei, et al.
Veröffentlicht: (2026)
PCA recovery thresholds in low-rank matrix inference with sparse noise
von: Adomaityte, Urte, et al.
Veröffentlicht: (2025)
von: Adomaityte, Urte, et al.
Veröffentlicht: (2025)
Efficient transformer adaptation for analog in-memory computing via low-rank adapters
von: Li, Chen, et al.
Veröffentlicht: (2024)
von: Li, Chen, et al.
Veröffentlicht: (2024)
On the Feature Learning in Diffusion Models
von: Han, Andi, et al.
Veröffentlicht: (2024)
von: Han, Andi, et al.
Veröffentlicht: (2024)
Harnessing small projectors and multiple views for efficient vision pretraining
von: Agrawal, Kumar Krishna, et al.
Veröffentlicht: (2023)
von: Agrawal, Kumar Krishna, et al.
Veröffentlicht: (2023)
Universal priors: solving empirical Bayes via Bayesian inference and pretraining
von: Cannella, Nick, et al.
Veröffentlicht: (2026)
von: Cannella, Nick, et al.
Veröffentlicht: (2026)
Evolution Strategies for Deep RL pretraining
von: Martínez, Adrian, et al.
Veröffentlicht: (2026)
von: Martínez, Adrian, et al.
Veröffentlicht: (2026)
Robust Least-Squares Optimization for Data-Driven Predictive Control: A Geometric Approach
von: Bharadwaj, Shreyas, et al.
Veröffentlicht: (2025)
von: Bharadwaj, Shreyas, et al.
Veröffentlicht: (2025)
The effectiveness of MAE pre-pretraining for billion-scale pretraining
von: Singh, Mannat, et al.
Veröffentlicht: (2023)
von: Singh, Mannat, et al.
Veröffentlicht: (2023)
On efficiently computable functions, deep networks and sparse compositionality
von: Poggio, Tomaso
Veröffentlicht: (2025)
von: Poggio, Tomaso
Veröffentlicht: (2025)
Storing overlapping associative memories on latent manifolds in low-rank spiking networks
von: Podlaski, William F., et al.
Veröffentlicht: (2024)
von: Podlaski, William F., et al.
Veröffentlicht: (2024)
Physics-informed waveform inversion using pretrained wavefield neural operators
von: Huang, Xinquan, et al.
Veröffentlicht: (2025)
von: Huang, Xinquan, et al.
Veröffentlicht: (2025)
Synthetic continued pretraining
von: Yang, Zitong, et al.
Veröffentlicht: (2024)
von: Yang, Zitong, et al.
Veröffentlicht: (2024)
MIRO: MultI-Reward cOnditioned pretraining improves T2I quality and efficiency
von: Dufour, Nicolas, et al.
Veröffentlicht: (2025)
von: Dufour, Nicolas, et al.
Veröffentlicht: (2025)
Convergence Error Analysis of Reflected Gradient Langevin Dynamics for Globally Optimizing Non-Convex Constrained Problems
von: Sato, Kanji, et al.
Veröffentlicht: (2022)
von: Sato, Kanji, et al.
Veröffentlicht: (2022)
A Tensor Residual Circuit Neural Network Factorized with Matrix Product Operation
von: Chen, Andi
Veröffentlicht: (2025)
von: Chen, Andi
Veröffentlicht: (2025)
Ähnliche Einträge
-
A Framework for Bilevel Optimization on Riemannian Manifolds
von: Han, Andi, et al.
Veröffentlicht: (2024) -
Riemannian coordinate descent algorithms on matrix manifolds
von: Han, Andi, et al.
Veröffentlicht: (2024) -
Generalized infinite dimensional Alpha-Procrustes based geometries
von: Goomanee, Salvish, et al.
Veröffentlicht: (2025) -
Riemannian Optimization for Hadamard Products of Low-Rank Matrices
von: Jawanpuria, Pratik, et al.
Veröffentlicht: (2026) -
A Gauss-Newton Approach for Min-Max Optimization in Generative Adversarial Networks
von: Mishra, Neel, et al.
Veröffentlicht: (2024)