Cosine-Gated Adam-Decay: Drop-In Staleness-Aware Outer Optimization for Decoupled DiLoCo
Fuente:
arXiv
Salvato in:
| Autori principali: | Shah, Vatsal, Sun, Jiahao |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
MuLoCo: Muon is a practical inner optimizer for DiLoCo
di: Thérien, Benjamin, et al.
Pubblicazione: (2025)
di: Thérien, Benjamin, et al.
Pubblicazione: (2025)
Smoothing DiLoCo with Primal Averaging for Faster Training of LLMs
di: Defazio, Aaron, et al.
Pubblicazione: (2025)
di: Defazio, Aaron, et al.
Pubblicazione: (2025)
DiLoCo: Distributed Low-Communication Training of Language Models
di: Douillard, Arthur, et al.
Pubblicazione: (2023)
di: Douillard, Arthur, et al.
Pubblicazione: (2023)
What happens when nanochat meets DiLoCo?
di: Acker, Alexander, et al.
Pubblicazione: (2025)
di: Acker, Alexander, et al.
Pubblicazione: (2025)
Communication-Efficient Language Model Training Scales Reliably and Robustly: Scaling Laws for DiLoCo
di: Charles, Zachary, et al.
Pubblicazione: (2025)
di: Charles, Zachary, et al.
Pubblicazione: (2025)
Decoupled DiLoCo for Resilient Distributed Pre-training
di: Douillard, Arthur, et al.
Pubblicazione: (2026)
di: Douillard, Arthur, et al.
Pubblicazione: (2026)
DP-AdamW: Investigating Decoupled Weight Decay and Bias Correction in Private Deep Learning
di: Chooi, Jay, et al.
Pubblicazione: (2025)
di: Chooi, Jay, et al.
Pubblicazione: (2025)
Eager Updates For Overlapped Communication and Computation in DiLoCo
di: Kale, Satyen, et al.
Pubblicazione: (2025)
di: Kale, Satyen, et al.
Pubblicazione: (2025)
LoCoCo: Dropping In Convolutions for Long Context Compression
di: Cai, Ruisi, et al.
Pubblicazione: (2024)
di: Cai, Ruisi, et al.
Pubblicazione: (2024)
LoCA: Location-Aware Cosine Adaptation for Parameter-Efficient Fine-Tuning
di: Du, Zhekai, et al.
Pubblicazione: (2025)
di: Du, Zhekai, et al.
Pubblicazione: (2025)
CosineGate: Semantic Dynamic Routing via Cosine Incompatibility in Residual Networks
di: Thota, Yogeswar Reddy
Pubblicazione: (2025)
di: Thota, Yogeswar Reddy
Pubblicazione: (2025)
Feature Staleness Aware Incremental Learning for CTR Prediction
di: Wang, Zhikai, et al.
Pubblicazione: (2025)
di: Wang, Zhikai, et al.
Pubblicazione: (2025)
Degree of Staleness-Aware Data Updating in Federated Learning
di: Liu, Tao, et al.
Pubblicazione: (2025)
di: Liu, Tao, et al.
Pubblicazione: (2025)
The Rich and the Simple: On the Implicit Bias of Adam and SGD
di: Vasudeva, Bhavya, et al.
Pubblicazione: (2025)
di: Vasudeva, Bhavya, et al.
Pubblicazione: (2025)
AdamHD: Decoupled Huber Decay Regularization for Language Model Pre-Training
di: Guo, Fu-Ming, et al.
Pubblicazione: (2025)
di: Guo, Fu-Ming, et al.
Pubblicazione: (2025)
VISAGNN: Versatile Staleness-Aware Efficient Training on Large-Scale Graphs
di: Xue, Rui
Pubblicazione: (2025)
di: Xue, Rui
Pubblicazione: (2025)
Beyond Cosine Decay: On the effectiveness of Infinite Learning Rate Schedule for Continual Pre-training
di: Singh, Vaibhav, et al.
Pubblicazione: (2025)
di: Singh, Vaibhav, et al.
Pubblicazione: (2025)
Correction of Decoupled Weight Decay
di: Chou, Jason Chuan-Chih
Pubblicazione: (2025)
di: Chou, Jason Chuan-Chih
Pubblicazione: (2025)
This Too Shall Pass: Removing Stale Observations in Dynamic Bayesian Optimization
di: Bardou, Anthony, et al.
Pubblicazione: (2024)
di: Bardou, Anthony, et al.
Pubblicazione: (2024)
Streaming DiLoCo with overlapping communication: Towards a Distributed Free Lunch
di: Douillard, Arthur, et al.
Pubblicazione: (2025)
di: Douillard, Arthur, et al.
Pubblicazione: (2025)
MSPipe: Efficient Temporal GNN Training via Staleness-Aware Pipeline
di: Sheng, Guangming, et al.
Pubblicazione: (2024)
di: Sheng, Guangming, et al.
Pubblicazione: (2024)
FedPSA: Modeling Behavioral Staleness in Asynchronous Federated Learning
di: Lu, Chaoyi, et al.
Pubblicazione: (2026)
di: Lu, Chaoyi, et al.
Pubblicazione: (2026)
Reviving Stale Updates: Data-Free Knowledge Distillation for Asynchronous Federated Learning
di: Askin, Baris, et al.
Pubblicazione: (2025)
di: Askin, Baris, et al.
Pubblicazione: (2025)
FedStaleWeight: Buffered Asynchronous Federated Learning with Fair Aggregation via Staleness Reweighting
di: Ma, Jeffrey, et al.
Pubblicazione: (2024)
di: Ma, Jeffrey, et al.
Pubblicazione: (2024)
Turning Stale Gradients into Stable Gradients: Coherent Coordinate Descent with Implicit Landscape Smoothing for Lightweight Zeroth-Order Optimization
di: Liang, Chen, et al.
Pubblicazione: (2026)
di: Liang, Chen, et al.
Pubblicazione: (2026)
FedLoDrop: Federated LoRA with Dropout for Generalized LLM Fine-tuning
di: Xie, Sijing, et al.
Pubblicazione: (2025)
di: Xie, Sijing, et al.
Pubblicazione: (2025)
Understanding Quantization of Optimizer States in LLM Pre-training: Dynamics of State Staleness and Effectiveness of State Resets
di: Topollai, Kristi, et al.
Pubblicazione: (2026)
di: Topollai, Kristi, et al.
Pubblicazione: (2026)
SoK: Blockchain-Based Decentralized AI (DeAI)
di: Lui, Elizabeth, et al.
Pubblicazione: (2024)
di: Lui, Elizabeth, et al.
Pubblicazione: (2024)
Cottention: Linear Transformers With Cosine Attention
di: Mongaras, Gabriel, et al.
Pubblicazione: (2024)
di: Mongaras, Gabriel, et al.
Pubblicazione: (2024)
The Hidden Pitfalls of the Cosine Similarity Loss
di: Draganov, Andrew, et al.
Pubblicazione: (2024)
di: Draganov, Andrew, et al.
Pubblicazione: (2024)
Distributed Perceptron under Bounded Staleness, Partial Participation, and Noisy Communication
di: Jain, Keval, et al.
Pubblicazione: (2026)
di: Jain, Keval, et al.
Pubblicazione: (2026)
Gated-SwinRMT: Unifying Swin Windowed Attention with Retentive Manhattan Decay via Input-Dependent Gating
di: Maity, Dipan, et al.
Pubblicazione: (2026)
di: Maity, Dipan, et al.
Pubblicazione: (2026)
Variance-Adjusted Cosine Distance as Similarity Metric
di: Sahoo, Satyajeet, et al.
Pubblicazione: (2025)
di: Sahoo, Satyajeet, et al.
Pubblicazione: (2025)
Accelerating Recommender Model Training by Dynamically Skipping Stale Embeddings
di: Maboud, Yassaman Ebrahimzadeh, et al.
Pubblicazione: (2024)
di: Maboud, Yassaman Ebrahimzadeh, et al.
Pubblicazione: (2024)
FedStale: leveraging stale client updates in federated learning
di: Rodio, Angelo, et al.
Pubblicazione: (2024)
di: Rodio, Angelo, et al.
Pubblicazione: (2024)
OpenDiLoCo: An Open-Source Framework for Globally Distributed Low-Communication Training
di: Jaghouar, Sami, et al.
Pubblicazione: (2024)
di: Jaghouar, Sami, et al.
Pubblicazione: (2024)
Word2VecGD: Neural Graph Drawing with Cosine-Stress Optimization
di: Yang, Minglai, et al.
Pubblicazione: (2025)
di: Yang, Minglai, et al.
Pubblicazione: (2025)
From Linear to Spline-Based Classification:Developing and Enhancing SMPA for Noisy Non-Linear Datasets
di: Srivastava, Vatsal
Pubblicazione: (2025)
di: Srivastava, Vatsal
Pubblicazione: (2025)
Understanding Adam Optimizer via Online Learning of Updates: Adam is FTRL in Disguise
di: Ahn, Kwangjun, et al.
Pubblicazione: (2024)
di: Ahn, Kwangjun, et al.
Pubblicazione: (2024)
Outer-Momentum Restarting in High-Dimensional Two-Phase Optimization
di: Topollai, Kristi, et al.
Pubblicazione: (2026)
di: Topollai, Kristi, et al.
Pubblicazione: (2026)
Documenti analoghi
-
MuLoCo: Muon is a practical inner optimizer for DiLoCo
di: Thérien, Benjamin, et al.
Pubblicazione: (2025) -
Smoothing DiLoCo with Primal Averaging for Faster Training of LLMs
di: Defazio, Aaron, et al.
Pubblicazione: (2025) -
DiLoCo: Distributed Low-Communication Training of Language Models
di: Douillard, Arthur, et al.
Pubblicazione: (2023) -
What happens when nanochat meets DiLoCo?
di: Acker, Alexander, et al.
Pubblicazione: (2025) -
Communication-Efficient Language Model Training Scales Reliably and Robustly: Scaling Laws for DiLoCo
di: Charles, Zachary, et al.
Pubblicazione: (2025)