OptiReduce: Resilient and Tail-Optimal AllReduce for Distributed Deep Learning in the Cloud
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Warraich, Ertza, Shabtai, Omer, Manaa, Khalid, Vargaftik, Shay, Piasetzky, Yonatan, Kadosh, Matty, Suresh, Lalith, Shahbaz, Muhammad |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
OptiNIC: A Resilient and Tail-Optimal RDMA NIC for Distributed ML Workloads
von: Warraich, Ertza, et al.
Veröffentlicht: (2025)
von: Warraich, Ertza, et al.
Veröffentlicht: (2025)
Reimagining RDMA Through the Lens of ML
von: Warraich, Ertza, et al.
Veröffentlicht: (2025)
von: Warraich, Ertza, et al.
Veröffentlicht: (2025)
Efficient AllReduce with Stragglers
von: Devraj, Arjun, et al.
Veröffentlicht: (2025)
von: Devraj, Arjun, et al.
Veröffentlicht: (2025)
Revisiting the Time Cost Model of AllReduce
von: Xiong, Dian, et al.
Veröffentlicht: (2024)
von: Xiong, Dian, et al.
Veröffentlicht: (2024)
Trivance: Latency-Optimal AllReduce by Shortcutting Multiport Networks
von: Juerss, Anton, et al.
Veröffentlicht: (2026)
von: Juerss, Anton, et al.
Veröffentlicht: (2026)
AllReduce Scheduling with Hierarchical Deep Reinforcement Learning
von: Wei, Yufan, et al.
Veröffentlicht: (2025)
von: Wei, Yufan, et al.
Veröffentlicht: (2025)
Short-circuiting Rings for Low-Latency AllReduce
von: Hammer, Sarah-Michelle, et al.
Veröffentlicht: (2025)
von: Hammer, Sarah-Michelle, et al.
Veröffentlicht: (2025)
Don't Let a Few Network Failures Slow the Entire AllReduce
von: Chen, Peiqing, et al.
Veröffentlicht: (2026)
von: Chen, Peiqing, et al.
Veröffentlicht: (2026)
Rina: Enhancing Ring-AllReduce with In-network Aggregation in Distributed Model Training
von: Chen, Zixuan, et al.
Veröffentlicht: (2024)
von: Chen, Zixuan, et al.
Veröffentlicht: (2024)
High-speed Networking for Giga-Scale AI Factories
von: Khashab, Sajy, et al.
Veröffentlicht: (2026)
von: Khashab, Sajy, et al.
Veröffentlicht: (2026)
VSS Challenge Problem: Verifying the Correctness of AllReduce Algorithms in the MPICH Implementation of MPI
von: Hovland, Paul D.
Veröffentlicht: (2025)
von: Hovland, Paul D.
Veröffentlicht: (2025)
DynamiQ: Accelerating Gradient Synchronization using Compressed Multi-hop All-reduce
von: Han, Wenchen, et al.
Veröffentlicht: (2026)
von: Han, Wenchen, et al.
Veröffentlicht: (2026)
Reducing Tail Latencies Through Environment- and Neighbour-aware Thread Management
von: Jeffery, Andrew, et al.
Veröffentlicht: (2024)
von: Jeffery, Andrew, et al.
Veröffentlicht: (2024)
HACK: Homomorphic Acceleration via Compression of the Key-Value Cache for Disaggregated LLM Inference
von: Zhang, Zeyu, et al.
Veröffentlicht: (2025)
von: Zhang, Zeyu, et al.
Veröffentlicht: (2025)
Optimal, Non-pipelined Reduce-scatter and Allreduce Algorithms
von: Träff, Jesper Larsson
Veröffentlicht: (2024)
von: Träff, Jesper Larsson
Veröffentlicht: (2024)
Near-Optimal Wafer-Scale Reduce
von: Luczynski, Piotr, et al.
Veröffentlicht: (2024)
von: Luczynski, Piotr, et al.
Veröffentlicht: (2024)
OptiLog: Assigning Roles in Byzantine Consensus
von: Gogada, Hanish, et al.
Veröffentlicht: (2025)
von: Gogada, Hanish, et al.
Veröffentlicht: (2025)
Spatio-Temporal Shifting to Reduce Carbon, Water, and Land-Use Footprints of Cloud Workloads
von: Attenni, Giulio, et al.
Veröffentlicht: (2025)
von: Attenni, Giulio, et al.
Veröffentlicht: (2025)
An All-Reduce Compatible Top-K Compressor for Communication-Efficient Distributed Learning
von: Chen, Chuyan, et al.
Veröffentlicht: (2025)
von: Chen, Chuyan, et al.
Veröffentlicht: (2025)
Õptimal Fault-Tolerant Labeling for Reachability and Approximate Distances in Directed Planar Graphs
von: Boneh, Itai, et al.
Veröffentlicht: (2025)
von: Boneh, Itai, et al.
Veröffentlicht: (2025)
Distributed Speculative Execution for Resilient Cloud Applications
von: Li, Tianyu, et al.
Veröffentlicht: (2024)
von: Li, Tianyu, et al.
Veröffentlicht: (2024)
Blockchain based Secure Energy Marketplace Scheme to Motivate Peer to Peer Microgrids
von: Awais, Muhammad, et al.
Veröffentlicht: (2022)
von: Awais, Muhammad, et al.
Veröffentlicht: (2022)
Low-ordered Orthogonal Voxel Finite Element with INT8 Tensor Cores for GPU-based Explicit Elastic Wave Propagation Analysis
von: Ichimura, Tsuyoshi, et al.
Veröffentlicht: (2024)
von: Ichimura, Tsuyoshi, et al.
Veröffentlicht: (2024)
Toward Optimal-Complexity Hash-Based Asynchronous MVBA with Optimal Resilience
von: Komatovic, Jovan, et al.
Veröffentlicht: (2024)
von: Komatovic, Jovan, et al.
Veröffentlicht: (2024)
Fault-tolerant Reduce and Allreduce operations based on correction
von: Kuettler, Martin, et al.
Veröffentlicht: (2026)
von: Kuettler, Martin, et al.
Veröffentlicht: (2026)
Resilience Evaluation of Kubernetes in Cloud-Edge Environments via Failure Injection
von: Chen, Zihao, et al.
Veröffentlicht: (2025)
von: Chen, Zihao, et al.
Veröffentlicht: (2025)
Amortized Asynchronous Byzantine Reliable Broadcast with Optimal Resilience
von: Hu, Michael Yiqing, et al.
Veröffentlicht: (2026)
von: Hu, Michael Yiqing, et al.
Veröffentlicht: (2026)
Meeting SLOs, Slashing Hours: Automated Enterprise LLM Optimization with OptiKIT
von: Santavas, Nicholas, et al.
Veröffentlicht: (2026)
von: Santavas, Nicholas, et al.
Veröffentlicht: (2026)
Accelerating Nonlinear Time-History Analysis with Complex Constitutive Laws via Heterogeneous Memory Management: From 3D Seismic Simulation to Neural Network Training
von: Ichimura, Tsuyoshi, et al.
Veröffentlicht: (2026)
von: Ichimura, Tsuyoshi, et al.
Veröffentlicht: (2026)
Failure-Resilient and Carbon-Efficient Deployment of Microservices over the Cloud-Edge Continuum
von: Ponce, Francisco, et al.
Veröffentlicht: (2026)
von: Ponce, Francisco, et al.
Veröffentlicht: (2026)
AgentFlow: Resilient Adaptive Cloud-Edge Framework for Multi-Agent Coordination
von: Chen, Ching Han, et al.
Veröffentlicht: (2025)
von: Chen, Ching Han, et al.
Veröffentlicht: (2025)
LA-IMR: Latency-Aware, Predictive In-Memory Routing and Proactive Autoscaling for Tail-Latency-Sensitive Cloud Robotics
von: Seo, Eunil, et al.
Veröffentlicht: (2025)
von: Seo, Eunil, et al.
Veröffentlicht: (2025)
Building Castles in the Cloud: Architecting Resilient and Scalable Infrastructure
von: Gundla, Naresh Kumar
Veröffentlicht: (2024)
von: Gundla, Naresh Kumar
Veröffentlicht: (2024)
Round and Resilience-Optimal Approximate Agreement on Trees and Block Graphs
von: Fuchs, Marc, et al.
Veröffentlicht: (2025)
von: Fuchs, Marc, et al.
Veröffentlicht: (2025)
CD-Raft: Reducing the Latency of Distributed Consensus in Cross-Domain Sites
von: Wang, Yangyang, et al.
Veröffentlicht: (2026)
von: Wang, Yangyang, et al.
Veröffentlicht: (2026)
AMSP: Reducing Communication Overhead of ZeRO for Efficient LLM Training
von: Chen, Qiaoling, et al.
Veröffentlicht: (2023)
von: Chen, Qiaoling, et al.
Veröffentlicht: (2023)
Reducing Energy Bloat in Large Model Training
von: Chung, Jae-Won, et al.
Veröffentlicht: (2023)
von: Chung, Jae-Won, et al.
Veröffentlicht: (2023)
UniPar: A Unified LLM-Based Framework for Parallel and Accelerated Code Translation in HPC
von: Bitan, Tomer, et al.
Veröffentlicht: (2025)
von: Bitan, Tomer, et al.
Veröffentlicht: (2025)
FuxiShuffle: An Adaptive and Resilient Shuffle Service for Distributed Data Processing on Alibaba Cloud
von: Lin, Yuhao, et al.
Veröffentlicht: (2026)
von: Lin, Yuhao, et al.
Veröffentlicht: (2026)
Scalable Fault-Tolerant MapReduce
von: Hespe, Demian, et al.
Veröffentlicht: (2024)
von: Hespe, Demian, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
OptiNIC: A Resilient and Tail-Optimal RDMA NIC for Distributed ML Workloads
von: Warraich, Ertza, et al.
Veröffentlicht: (2025) -
Reimagining RDMA Through the Lens of ML
von: Warraich, Ertza, et al.
Veröffentlicht: (2025) -
Efficient AllReduce with Stragglers
von: Devraj, Arjun, et al.
Veröffentlicht: (2025) -
Revisiting the Time Cost Model of AllReduce
von: Xiong, Dian, et al.
Veröffentlicht: (2024) -
Trivance: Latency-Optimal AllReduce by Shortcutting Multiport Networks
von: Juerss, Anton, et al.
Veröffentlicht: (2026)