MT-DAO: Multi-Timescale Distributed Adaptive Optimizers with Local Updates
Fuente:
arXiv
Saved in:
| Main Authors: | Iacob, Alex, Jovanovic, Andrej, Safaryan, Mher, Kurmanji, Meghdad, Sani, Lorenzo, Horváth, Samuel, Shen, William F., Qiu, Xinchi, Lane, Nicholas D. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DES-LOC: Desynced Low Communication Adaptive Optimizers for Training Foundation Models
by: Iacob, Alex, et al.
Published: (2025)
by: Iacob, Alex, et al.
Published: (2025)
LoRDO: Distributed Low-Rank Optimization with Infrequent Communication
by: Jovanović, Andrej, et al.
Published: (2026)
by: Jovanović, Andrej, et al.
Published: (2026)
Response-Conditioned Parallel-to-Sequential Orchestration for Multi-Agent Systems
by: Tastan, Nurbek, et al.
Published: (2026)
by: Tastan, Nurbek, et al.
Published: (2026)
DEPT: Decoupled Embeddings for Pre-training Language Models
by: Iacob, Alex, et al.
Published: (2024)
by: Iacob, Alex, et al.
Published: (2024)
LLM Unlearning via Neural Activation Redirection
by: Shen, William F., et al.
Published: (2025)
by: Shen, William F., et al.
Published: (2025)
AbbIE: Autoregressive Block-Based Iterative Encoder for Efficient Sequence Modeling
by: Aleksandrov, Preslav, et al.
Published: (2025)
by: Aleksandrov, Preslav, et al.
Published: (2025)
On Biased Compression for Distributed Learning
by: Beznosikov, Aleksandr, et al.
Published: (2020)
by: Beznosikov, Aleksandr, et al.
Published: (2020)
Position: Bridge the Gaps between Machine Unlearning and AI Regulation
by: Marino, Bill, et al.
Published: (2025)
by: Marino, Bill, et al.
Published: (2025)
How Data Inter-connectivity Shapes LLMs Unlearning: A Structural Unlearning Perspective
by: Qiu, Xinchi, et al.
Published: (2024)
by: Qiu, Xinchi, et al.
Published: (2024)
Editing as Unlearning: Are Knowledge Editing Methods Strong Baselines for Large Language Model Unlearning?
by: Li, Zexi, et al.
Published: (2025)
by: Li, Zexi, et al.
Published: (2025)
GradSkip: Communication-Accelerated Local Gradient Methods with Better Computational Complexity
by: Maranjyan, Artavazd, et al.
Published: (2022)
by: Maranjyan, Artavazd, et al.
Published: (2022)
SparsyFed: Sparse Adaptive Federated Training
by: Guastella, Adriano, et al.
Published: (2025)
by: Guastella, Adriano, et al.
Published: (2025)
LDAdam: Adaptive Optimization from Low-Dimensional Gradient Statistics
by: Robert, Thomas, et al.
Published: (2024)
by: Robert, Thomas, et al.
Published: (2024)
Sheaf HyperNetworks for Personalized Federated Learning
by: Nguyen, Bao, et al.
Published: (2024)
by: Nguyen, Bao, et al.
Published: (2024)
$f$-FUM: Federated Unlearning via min--max and $f$-divergence
by: Karimian, Radmehr, et al.
Published: (2026)
by: Karimian, Radmehr, et al.
Published: (2026)
Pollen: High-throughput Federated Learning Simulation via Resource-Aware Client Placement
by: Sani, Lorenzo, et al.
Published: (2023)
by: Sani, Lorenzo, et al.
Published: (2023)
FedAnchor: Enhancing Federated Semi-Supervised Learning with Label Contrastive Loss for Unlabeled Clients
by: Qiu, Xinchi, et al.
Published: (2024)
by: Qiu, Xinchi, et al.
Published: (2024)
Worldwide Federated Training of Language Models
by: Iacob, Alex, et al.
Published: (2024)
by: Iacob, Alex, et al.
Published: (2024)
The Future of Large Language Model Pre-training is Federated
by: Sani, Lorenzo, et al.
Published: (2024)
by: Sani, Lorenzo, et al.
Published: (2024)
Towards Robust Scaling Laws for Optimizers
by: Volkova, Alexandra, et al.
Published: (2026)
by: Volkova, Alexandra, et al.
Published: (2026)
Photon: Federated LLM Pre-Training
by: Sani, Lorenzo, et al.
Published: (2024)
by: Sani, Lorenzo, et al.
Published: (2024)
MatryoshkaLoRA: Learning Accurate Hierarchical Low-Rank Representations for LLM Fine-Tuning
by: Modoranu, Ionut-Vlad, et al.
Published: (2026)
by: Modoranu, Ionut-Vlad, et al.
Published: (2026)
FFT-based Dynamic Subspace Selection for Low-Rank Adaptive Optimization of Large Language Models
by: Modoranu, Ionut-Vlad, et al.
Published: (2025)
by: Modoranu, Ionut-Vlad, et al.
Published: (2025)
Task-Centric Personalized Federated Fine-Tuning of Language Models
by: Talasso, Gabriel U., et al.
Published: (2026)
by: Talasso, Gabriel U., et al.
Published: (2026)
MicroAdam: Accurate Adaptive Optimization with Low Space Overhead and Provable Convergence
by: Modoranu, Ionut-Vlad, et al.
Published: (2024)
by: Modoranu, Ionut-Vlad, et al.
Published: (2024)
SEAT: Sparse Entity-Aware Tuning for Knowledge Adaptation while Preserving Epistemic Abstention
by: Shen, William F., et al.
Published: (2025)
by: Shen, William F., et al.
Published: (2025)
CAGE: Curvature-Aware Gradient Estimation For Accurate Quantization-Aware Training
by: Tabesh, Soroush, et al.
Published: (2025)
by: Tabesh, Soroush, et al.
Published: (2025)
The Iterative Optimal Brain Surgeon: Faster Sparse Recovery by Leveraging Second-Order Information
by: Wu, Diyuan, et al.
Published: (2024)
by: Wu, Diyuan, et al.
Published: (2024)
DASH: Faster Shampoo via Batched Block Preconditioning and Efficient Inverse-Root Solvers
by: Modoranu, Ionut-Vlad, et al.
Published: (2026)
by: Modoranu, Ionut-Vlad, et al.
Published: (2026)
What makes unlearning hard and what to do about it
by: Zhao, Kairan, et al.
Published: (2024)
by: Zhao, Kairan, et al.
Published: (2024)
Convergence of Distributed Adaptive Optimization with Local Updates
by: Cheng, Ziheng, et al.
Published: (2024)
by: Cheng, Ziheng, et al.
Published: (2024)
Gradient-less Federated Gradient Boosting Trees with Learnable Learning Rates
by: Ma, Chenyang, et al.
Published: (2023)
by: Ma, Chenyang, et al.
Published: (2023)
Public Persuasion with Endogenous Fact-Checking
by: Lukyanov, Georgy, et al.
Published: (2025)
by: Lukyanov, Georgy, et al.
Published: (2025)
Unified Scaling Laws for Compressed Representations
by: Panferov, Andrei, et al.
Published: (2025)
by: Panferov, Andrei, et al.
Published: (2025)
DAO to (Anonymous) DAO Transactions
by: Qi, Minfeng, et al.
Published: (2026)
by: Qi, Minfeng, et al.
Published: (2026)
Hymnographic Indicators of the Armenian Renaissance
by: Mher Navoyan
Published: (2026)
by: Mher Navoyan
Published: (2026)
The On-Chain and Off-Chain Mechanisms of DAO-to-DAO Voting
by: Lloyd, Thomas, et al.
Published: (2026)
by: Lloyd, Thomas, et al.
Published: (2026)
Improving the ROS 2 Navigation Stack with Real-Time Local Costmap Updates for Agricultural Applications
by: Sani, Ettore, et al.
Published: (2024)
by: Sani, Ettore, et al.
Published: (2024)
AgentDAO: Synthesis of Proposal Transactions Via Abstract DAO Semantics
by: Ao, Lin, et al.
Published: (2025)
by: Ao, Lin, et al.
Published: (2025)
Tableau Proof Systems for Justification Logics
by: Ghari, Meghdad
Published: (2014)
by: Ghari, Meghdad
Published: (2014)
Similar Items
-
DES-LOC: Desynced Low Communication Adaptive Optimizers for Training Foundation Models
by: Iacob, Alex, et al.
Published: (2025) -
LoRDO: Distributed Low-Rank Optimization with Infrequent Communication
by: Jovanović, Andrej, et al.
Published: (2026) -
Response-Conditioned Parallel-to-Sequential Orchestration for Multi-Agent Systems
by: Tastan, Nurbek, et al.
Published: (2026) -
DEPT: Decoupled Embeddings for Pre-training Language Models
by: Iacob, Alex, et al.
Published: (2024) -
LLM Unlearning via Neural Activation Redirection
by: Shen, William F., et al.
Published: (2025)