LoCo: Low-Bit Communication Adaptor for Large-scale Model Training
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xie, Xingyu, Lin, Zhijie, Toh, Kim-Chuan, Zhou, Pan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Optimization Hyper-parameter Laws for Large Language Models
von: Xie, Xingyu, et al.
Veröffentlicht: (2024)
von: Xie, Xingyu, et al.
Veröffentlicht: (2024)
On exploration of an interior mirror descent flow for stochastic nonconvex constrained problem
von: Ding, Kuangyu, et al.
Veröffentlicht: (2025)
von: Ding, Kuangyu, et al.
Veröffentlicht: (2025)
Nonconvex Stochastic Bregman Proximal Gradient Method with Application to Deep Learning
von: Ding, Kuangyu, et al.
Veröffentlicht: (2023)
von: Ding, Kuangyu, et al.
Veröffentlicht: (2023)
Learning Graph Laplacian with MCP
von: Zhang, Yangjing, et al.
Veröffentlicht: (2020)
von: Zhang, Yangjing, et al.
Veröffentlicht: (2020)
Adan: Adaptive Nesterov Momentum Algorithm for Faster Optimizing Deep Models
von: Xie, Xingyu, et al.
Veröffentlicht: (2022)
von: Xie, Xingyu, et al.
Veröffentlicht: (2022)
Developing Lagrangian-based Methods for Nonsmooth Nonconvex Optimization
von: Xiao, Nachuan, et al.
Veröffentlicht: (2024)
von: Xiao, Nachuan, et al.
Veröffentlicht: (2024)
Adam-family Methods for Nonsmooth Optimization with Convergence Guarantees
von: Xiao, Nachuan, et al.
Veröffentlicht: (2023)
von: Xiao, Nachuan, et al.
Veröffentlicht: (2023)
BiCoLoR: Communication-Efficient Optimization with Bidirectional Compression and Local Training
von: Condat, Laurent, et al.
Veröffentlicht: (2026)
von: Condat, Laurent, et al.
Veröffentlicht: (2026)
Memory-Efficient 4-bit Preconditioned Stochastic Optimization
von: Li, Jingyang, et al.
Veröffentlicht: (2024)
von: Li, Jingyang, et al.
Veröffentlicht: (2024)
Accelerating nuclear-norm regularized low-rank matrix optimization through Burer-Monteiro decomposition
von: Lee, Ching-pei, et al.
Veröffentlicht: (2022)
von: Lee, Ching-pei, et al.
Veröffentlicht: (2022)
Stochastic Subgradient Methods with Guaranteed Global Stability in Nonsmooth Nonconvex Optimization
von: Xiao, Nachuan, et al.
Veröffentlicht: (2023)
von: Xiao, Nachuan, et al.
Veröffentlicht: (2023)
GNMR: Runtime Stability Control for Low-Precision Large Language Model Training
von: Kong, Boao, et al.
Veröffentlicht: (2026)
von: Kong, Boao, et al.
Veröffentlicht: (2026)
Tractable hierarchies of convex relaxations for polynomial optimization on the nonnegative orthant
von: Mai, Ngoc Hoang Anh, et al.
Veröffentlicht: (2022)
von: Mai, Ngoc Hoang Anh, et al.
Veröffentlicht: (2022)
Multi-Objective Linear Ensembles for Robust and Sparse Training of Few-Bit Neural Networks
von: Bernardelli, Ambrogio Maria, et al.
Veröffentlicht: (2022)
von: Bernardelli, Ambrogio Maria, et al.
Veröffentlicht: (2022)
LoCoDL: Communication-Efficient Distributed Learning with Local Training and Compression
von: Condat, Laurent, et al.
Veröffentlicht: (2024)
von: Condat, Laurent, et al.
Veröffentlicht: (2024)
MARS: Unleashing the Power of Variance Reduction for Training Large Models
von: Yuan, Huizhuo, et al.
Veröffentlicht: (2024)
von: Yuan, Huizhuo, et al.
Veröffentlicht: (2024)
Low-rank Orthogonalization for Large-scale Matrix Optimization with Applications to Foundation Model Training
von: He, Chuan, et al.
Veröffentlicht: (2025)
von: He, Chuan, et al.
Veröffentlicht: (2025)
LoRA Training in the NTK Regime has No Spurious Local Minima
von: Jang, Uijeong, et al.
Veröffentlicht: (2024)
von: Jang, Uijeong, et al.
Veröffentlicht: (2024)
AdLoCo: adaptive batching significantly improves communications efficiency and convergence for Large Language Models
von: Kutuzov, Nikolay, et al.
Veröffentlicht: (2025)
von: Kutuzov, Nikolay, et al.
Veröffentlicht: (2025)
LoFT: Low-Rank Adaptation That Behaves Like Full Fine-Tuning
von: Tastan, Nurbek, et al.
Veröffentlicht: (2025)
von: Tastan, Nurbek, et al.
Veröffentlicht: (2025)
Bernoulli-LoRA: A Theoretical Framework for Randomized Low-Rank Adaptation
von: Sokolov, Igor, et al.
Veröffentlicht: (2025)
von: Sokolov, Igor, et al.
Veröffentlicht: (2025)
Randomized Asymmetric Chain of LoRA: The First Meaningful Theoretical Framework for Low-Rank Adaptation
von: Malinovsky, Grigory, et al.
Veröffentlicht: (2024)
von: Malinovsky, Grigory, et al.
Veröffentlicht: (2024)
Training Deep Learning Models with Norm-Constrained LMOs
von: Pethick, Thomas, et al.
Veröffentlicht: (2025)
von: Pethick, Thomas, et al.
Veröffentlicht: (2025)
An Overview of Low-Rank Structures in the Training and Adaptation of Large Models
von: Balzano, Laura, et al.
Veröffentlicht: (2025)
von: Balzano, Laura, et al.
Veröffentlicht: (2025)
Gathering and Exploiting Higher-Order Information when Training Large Structured Models
von: Wolinski, Pierre
Veröffentlicht: (2023)
von: Wolinski, Pierre
Veröffentlicht: (2023)
DOVA-PATBM: An Intelligent, Adaptive, and Scalable Framework for Optimizing Large-Scale EV Charging Infrastructure
von: Li, Chuan, et al.
Veröffentlicht: (2025)
von: Li, Chuan, et al.
Veröffentlicht: (2025)
Optimal Rates for Robust Stochastic Convex Optimization
von: Gao, Changyu, et al.
Veröffentlicht: (2024)
von: Gao, Changyu, et al.
Veröffentlicht: (2024)
On the B-subdifferential of proximal operators of affine-constrained $\ell_1$ regularizer
von: Li, Xudong, et al.
Veröffentlicht: (2025)
von: Li, Xudong, et al.
Veröffentlicht: (2025)
The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training
von: Schaipp, Fabian, et al.
Veröffentlicht: (2025)
von: Schaipp, Fabian, et al.
Veröffentlicht: (2025)
Stochastic Bregman Subgradient Methods for Nonsmooth Nonconvex Optimization Problems
von: Ding, Kuangyu, et al.
Veröffentlicht: (2024)
von: Ding, Kuangyu, et al.
Veröffentlicht: (2024)
Inexact Bregman Proximal Gradient Method and its Inertial Variant with Absolute and Partial Relative Stopping Criteria
von: Yang, Lei, et al.
Veröffentlicht: (2021)
von: Yang, Lei, et al.
Veröffentlicht: (2021)
StoSignSGD: Unbiased Structural Stochasticity Fixes SignSGD for Training Large Language Models
von: Yu, Dingzhi, et al.
Veröffentlicht: (2026)
von: Yu, Dingzhi, et al.
Veröffentlicht: (2026)
Stable Nonconvex-Nonconcave Training via Linear Interpolation
von: Pethick, Thomas, et al.
Veröffentlicht: (2023)
von: Pethick, Thomas, et al.
Veröffentlicht: (2023)
Private Heterogeneous Federated Learning Without a Trusted Server Revisited: Error-Optimal and Communication-Efficient Algorithms for Convex Losses
von: Gao, Changyu, et al.
Veröffentlicht: (2024)
von: Gao, Changyu, et al.
Veröffentlicht: (2024)
NewVEM: A Newton Vertex Exchange Method for a Class of Constrained Self-Concordant Minimization Problems
von: Liang, Ling, et al.
Veröffentlicht: (2024)
von: Liang, Ling, et al.
Veröffentlicht: (2024)
The Power of Preconditioning in Overparameterized Low-Rank Matrix Sensing
von: Xu, Xingyu, et al.
Veröffentlicht: (2023)
von: Xu, Xingyu, et al.
Veröffentlicht: (2023)
Stability and Generalization for Stochastic Recursive Momentum-based Algorithms for (Strongly-)Convex One to $K$-Level Stochastic Optimizations
von: Pan, Xiaokang, et al.
Veröffentlicht: (2024)
von: Pan, Xiaokang, et al.
Veröffentlicht: (2024)
NeuralQP: A General Hypergraph-based Optimization Framework for Large-scale QCQPs
von: Xiong, Zhixiao, et al.
Veröffentlicht: (2024)
von: Xiong, Zhixiao, et al.
Veröffentlicht: (2024)
Wasserstein distributionally robust optimization and its tractable regularization formulations
von: Chu, Hong T. M., et al.
Veröffentlicht: (2024)
von: Chu, Hong T. M., et al.
Veröffentlicht: (2024)
From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees
von: Xie, Shengping, et al.
Veröffentlicht: (2025)
von: Xie, Shengping, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Optimization Hyper-parameter Laws for Large Language Models
von: Xie, Xingyu, et al.
Veröffentlicht: (2024) -
On exploration of an interior mirror descent flow for stochastic nonconvex constrained problem
von: Ding, Kuangyu, et al.
Veröffentlicht: (2025) -
Nonconvex Stochastic Bregman Proximal Gradient Method with Application to Deep Learning
von: Ding, Kuangyu, et al.
Veröffentlicht: (2023) -
Learning Graph Laplacian with MCP
von: Zhang, Yangjing, et al.
Veröffentlicht: (2020) -
Adan: Adaptive Nesterov Momentum Algorithm for Faster Optimizing Deep Models
von: Xie, Xingyu, et al.
Veröffentlicht: (2022)