Fault Tolerant ML: Efficient Meta-Aggregation and Synchronous Training
Fuente:
arXiv
Saved in:
| Main Authors: | Dahan, Tehila, Levy, Kfir Y. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Weight for Robustness: A Comprehensive Approach towards Optimal Fault-Tolerant Asynchronous ML
by: Dahan, Tehila, et al.
Published: (2025)
by: Dahan, Tehila, et al.
Published: (2025)
SLowcal-SGD: Slow Query Points Improve Local-SGD for Stochastic Convex Optimization
by: Dahan, Tehila, et al.
Published: (2023)
by: Dahan, Tehila, et al.
Published: (2023)
$μ^2$-SGD: Stable Stochastic Optimization via a Double Momentum Mechanism
by: Dahan, Tehila, et al.
Published: (2023)
by: Dahan, Tehila, et al.
Published: (2023)
Bringing Order to Asynchronous SGD: Towards Optimality under Data-Dependent Delays with Momentum
by: Dahan, Tehila, et al.
Published: (2026)
by: Dahan, Tehila, et al.
Published: (2026)
Local MixVR: Breaking the Communication-Sample Dependence in Distributed Learning
by: Dahan, Tehila, et al.
Published: (2026)
by: Dahan, Tehila, et al.
Published: (2026)
Optimal Sample Complexity for Single Time-Scale Actor-Critic with Momentum
by: Kumar, Navdeep, et al.
Published: (2026)
by: Kumar, Navdeep, et al.
Published: (2026)
Private and Federated Stochastic Convex Optimization: Efficient Strategies for Centralized Systems
by: Reshef, Roie, et al.
Published: (2024)
by: Reshef, Roie, et al.
Published: (2024)
Safety in the Face of Adversity: Achieving Zero Constraint Violation in Online Learning with Slowly Changing Constraints
by: Hamoud, Bassel, et al.
Published: (2025)
by: Hamoud, Bassel, et al.
Published: (2025)
Enhancing Parallelism in Decentralized Stochastic Convex Optimization
by: Eisen, Ofri, et al.
Published: (2025)
by: Eisen, Ofri, et al.
Published: (2025)
Privacy-Preserving Federated Convex Optimization: Balancing Partial-Participation and Efficiency via Noise Cancellation
by: Reshef, Roie, et al.
Published: (2025)
by: Reshef, Roie, et al.
Published: (2025)
Beyond Communication Overhead: A Multilevel Monte Carlo Approach for Mitigating Compression Bias in Distributed Learning
by: Zukerman, Ze'ev, et al.
Published: (2025)
by: Zukerman, Ze'ev, et al.
Published: (2025)
Dynamic Byzantine-Robust Learning: Adapting to Switching Byzantine Workers
by: Dorfman, Ron, et al.
Published: (2024)
by: Dorfman, Ron, et al.
Published: (2024)
Fault-Tolerant Evaluation for Sample-Efficient Model Performance Estimators
by: Zhu, Zihan, et al.
Published: (2026)
by: Zhu, Zihan, et al.
Published: (2026)
DE-PADA: Personalized Augmentation and Domain Adaptation for ECG Biometrics Across Physiological States
by: Saleh, Amro Abu, et al.
Published: (2025)
by: Saleh, Amro Abu, et al.
Published: (2025)
Prediction-Powered Semi-Supervised Learning with Online Power Tuning
by: Shoham, Noa, et al.
Published: (2025)
by: Shoham, Noa, et al.
Published: (2025)
MoC-System: Efficient Fault Tolerance for Sparse Mixture-of-Experts Model Training
by: Cai, Weilin, et al.
Published: (2024)
by: Cai, Weilin, et al.
Published: (2024)
Training a Bilingual Language Model by Mapping Tokens onto a Shared Character Space
by: Rom, Aviad, et al.
Published: (2024)
by: Rom, Aviad, et al.
Published: (2024)
Accelerating Distributed ML Training via Selective Synchronization
by: Tyagi, Sahil, et al.
Published: (2023)
by: Tyagi, Sahil, et al.
Published: (2023)
Bring Your Own (Non-Robust) Algorithm to Solve Robust MDPs by Estimating The Worst Kernel
by: Wang, Kaixin, et al.
Published: (2023)
by: Wang, Kaixin, et al.
Published: (2023)
The Streaming Batch Model for Efficient and Fault-Tolerant Heterogeneous Execution
by: Luan, Frank Sifei, et al.
Published: (2025)
by: Luan, Frank Sifei, et al.
Published: (2025)
Fed-Meta-Align: A Similarity-Aware Aggregation and Personalization Pipeline for Federated TinyML on Heterogeneous Data
by: Macharla, Hemanth, et al.
Published: (2025)
by: Macharla, Hemanth, et al.
Published: (2025)
On the Byzantine Fault Tolerance of signSGD with Majority Vote
by: Mengoli, Emanuele, et al.
Published: (2025)
by: Mengoli, Emanuele, et al.
Published: (2025)
On the Convergence of Single-Timescale Actor-Critic
by: Kumar, Navdeep, et al.
Published: (2024)
by: Kumar, Navdeep, et al.
Published: (2024)
On the Global Convergence of Policy Gradient in Average Reward Markov Decision Processes
by: Kumar, Navdeep, et al.
Published: (2024)
by: Kumar, Navdeep, et al.
Published: (2024)
Role-Based Fault Tolerance System for LLM RL Post-Training
by: Chen, Zhenqian, et al.
Published: (2025)
by: Chen, Zhenqian, et al.
Published: (2025)
Gradient-Variation Online Adaptivity for Accelerated Optimization with Hölder Smoothness
by: Zhao, Yuheng, et al.
Published: (2025)
by: Zhao, Yuheng, et al.
Published: (2025)
Embedding Byzantine Fault Tolerance into Federated Learning via Consistency Scoring
by: Lee, Youngjoon, et al.
Published: (2024)
by: Lee, Youngjoon, et al.
Published: (2024)
Efficient and Microphone-Fault-Tolerant 3D Sound Source Localization
by: Yang, Yiyuan, et al.
Published: (2025)
by: Yang, Yiyuan, et al.
Published: (2025)
Solving Non-Rectangular Reward-Robust MDPs via Frequency Regularization
by: Gadot, Uri, et al.
Published: (2023)
by: Gadot, Uri, et al.
Published: (2023)
FT-MoE: Sustainable-learning Mixture of Experts for Fault-Tolerant Computing
by: Xiao, Wenjing, et al.
Published: (2025)
by: Xiao, Wenjing, et al.
Published: (2025)
Training ML Models with Predictable Failures
by: Schwarzer, Will, et al.
Published: (2026)
by: Schwarzer, Will, et al.
Published: (2026)
Custom Algorithm-based Fault Tolerance for Attention Layers in Transformers
by: Titopoulos, Vasileios, et al.
Published: (2025)
by: Titopoulos, Vasileios, et al.
Published: (2025)
FORTALESA: Fault-Tolerant Reconfigurable Systolic Array for DNN Inference
by: Cherezova, Natalia, et al.
Published: (2025)
by: Cherezova, Natalia, et al.
Published: (2025)
Cost-Effective Fault Tolerance for CNNs Using Parameter Vulnerability Based Hardening and Pruning
by: Ahmadilivani, Mohammad Hasan, et al.
Published: (2024)
by: Ahmadilivani, Mohammad Hasan, et al.
Published: (2024)
Development of an Edge Resilient ML Ensemble to Tolerate ICS Adversarial Attacks
by: Yao, Likai, et al.
Published: (2024)
by: Yao, Likai, et al.
Published: (2024)
Bilevel ZOFO: Efficient LLM Fine-Tuning and Meta-Training
by: Shirkavand, Reza, et al.
Published: (2025)
by: Shirkavand, Reza, et al.
Published: (2025)
Towards Fault Tolerance in Multi-Agent Reinforcement Learning
by: Shi, Yuchen, et al.
Published: (2024)
by: Shi, Yuchen, et al.
Published: (2024)
Layer-Specific Lipschitz Modulation for Fault-Tolerant Multimodal Representation Learning
by: Altinses, Diyar, et al.
Published: (2026)
by: Altinses, Diyar, et al.
Published: (2026)
Efficient Training with Denoised Neural Weights
by: Gong, Yifan, et al.
Published: (2024)
by: Gong, Yifan, et al.
Published: (2024)
MetaML-Pro: Cross-Stage Design Flow Automation for Efficient Deep Learning Acceleration
by: Que, Zhiqiang, et al.
Published: (2025)
by: Que, Zhiqiang, et al.
Published: (2025)
Similar Items
-
Weight for Robustness: A Comprehensive Approach towards Optimal Fault-Tolerant Asynchronous ML
by: Dahan, Tehila, et al.
Published: (2025) -
SLowcal-SGD: Slow Query Points Improve Local-SGD for Stochastic Convex Optimization
by: Dahan, Tehila, et al.
Published: (2023) -
$μ^2$-SGD: Stable Stochastic Optimization via a Double Momentum Mechanism
by: Dahan, Tehila, et al.
Published: (2023) -
Bringing Order to Asynchronous SGD: Towards Optimality under Data-Dependent Delays with Momentum
by: Dahan, Tehila, et al.
Published: (2026) -
Local MixVR: Breaking the Communication-Sample Dependence in Distributed Learning
by: Dahan, Tehila, et al.
Published: (2026)