Chimera: Efficiently Training Large-Scale Neural Networks with Bidirectional Pipelines
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Shigang, Hoefler, Torsten |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2021
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Near-Optimal Sparse Allreduce for Distributed Deep Learning
von: Li, Shigang, et al.
Veröffentlicht: (2022)
von: Li, Shigang, et al.
Veröffentlicht: (2022)
AutoDDL: Automatic Distributed Deep Learning with Near-Optimal Bandwidth Cost
von: Chen, Jinfan, et al.
Veröffentlicht: (2023)
von: Chen, Jinfan, et al.
Veröffentlicht: (2023)
SparkAttention: High-Performance Multi-Head Attention for Large Models on Volta GPU Architecture
von: Xu, Youxuan, et al.
Veröffentlicht: (2025)
von: Xu, Youxuan, et al.
Veröffentlicht: (2025)
FlashSparse: Minimizing Computation Redundancy for Fast Sparse Matrix Multiplications on Tensor Cores
von: Shi, Jinliang, et al.
Veröffentlicht: (2024)
von: Shi, Jinliang, et al.
Veröffentlicht: (2024)
Libra: Unleashing GPU Heterogeneity for High-Performance Sparse Matrix Multiplication
von: Shi, Jinliang, et al.
Veröffentlicht: (2025)
von: Shi, Jinliang, et al.
Veröffentlicht: (2025)
How Machine Learning-Data Driven Replication Strategies Enhance Fault Tolerance in Large-Scale Distributed Systems
von: Murimi, Almond Kiruthu
Veröffentlicht: (2025)
von: Murimi, Almond Kiruthu
Veröffentlicht: (2025)
Impact of Network Topology on Byzantine Resilience in Decentralized Federated Learning
von: Bhattacharya, Siddhartha, et al.
Veröffentlicht: (2024)
von: Bhattacharya, Siddhartha, et al.
Veröffentlicht: (2024)
Token Coherence: Adapting MESI Cache Protocols to Minimize Synchronization Overhead in Multi-Agent LLM Systems
von: Parakhin, Vladyslav
Veröffentlicht: (2026)
von: Parakhin, Vladyslav
Veröffentlicht: (2026)
HFedATM: Hierarchical Federated Domain Generalization via Optimal Transport and Regularized Mean Aggregation
von: Nguyen, Thinh, et al.
Veröffentlicht: (2025)
von: Nguyen, Thinh, et al.
Veröffentlicht: (2025)
TAGC: Optimizing Gradient Communication in Distributed Transformer Training
von: Polyakov, Igor, et al.
Veröffentlicht: (2025)
von: Polyakov, Igor, et al.
Veröffentlicht: (2025)
AMP4EC: Adaptive Model Partitioning Framework for Efficient Deep Learning Inference in Edge Computing Environments
von: Zhang, Guilin, et al.
Veröffentlicht: (2025)
von: Zhang, Guilin, et al.
Veröffentlicht: (2025)
SAGA: Workflow-Atomic Scheduling for AI Agent Inference on GPU Clusters
von: Guo, Dongxin, et al.
Veröffentlicht: (2026)
von: Guo, Dongxin, et al.
Veröffentlicht: (2026)
LayerKV: Optimizing Large Language Model Serving with Layer-wise KV Cache Management
von: Xiong, Yi, et al.
Veröffentlicht: (2024)
von: Xiong, Yi, et al.
Veröffentlicht: (2024)
TokenCake: A KV-Cache-centric Serving Framework for LLM-based Multi-Agent Applications
von: Bian, Zhuohang, et al.
Veröffentlicht: (2025)
von: Bian, Zhuohang, et al.
Veröffentlicht: (2025)
GRAIN: Exact Graph Reconstruction from Gradients
von: Drencheva, Maria, et al.
Veröffentlicht: (2025)
von: Drencheva, Maria, et al.
Veröffentlicht: (2025)
Accelerating Geo-distributed Machine Learning with Network-Aware Adaptive Tree and Auxiliary Route
von: Li, Zonghang, et al.
Veröffentlicht: (2024)
von: Li, Zonghang, et al.
Veröffentlicht: (2024)
Federated Learning Model Aggregation in Heterogenous Aerial and Space Networks
von: Dong, Fan, et al.
Veröffentlicht: (2023)
von: Dong, Fan, et al.
Veröffentlicht: (2023)
Training Diffusion Models with Federated Learning
von: de Goede, Matthijs, et al.
Veröffentlicht: (2024)
von: de Goede, Matthijs, et al.
Veröffentlicht: (2024)
Training LLMs on HPC Systems: Best Practices from the OpenGPT-X Project
von: Penke, Carolin, et al.
Veröffentlicht: (2025)
von: Penke, Carolin, et al.
Veröffentlicht: (2025)
Quantize Once, Train Fast: Allreduce-Compatible Compression with Provable Guarantees
von: Xin, Jihao, et al.
Veröffentlicht: (2023)
von: Xin, Jihao, et al.
Veröffentlicht: (2023)
A Survey on Efficient Federated Learning Methods for Foundation Model Training
von: Woisetschläger, Herbert, et al.
Veröffentlicht: (2024)
von: Woisetschläger, Herbert, et al.
Veröffentlicht: (2024)
SI-ChainFL: Shapley-Incentivized Secure Federated Learning for High-Speed Rail Data Sharing
von: Zhao, Mingjie, et al.
Veröffentlicht: (2026)
von: Zhao, Mingjie, et al.
Veröffentlicht: (2026)
Breaking (Global) Barriers in Parallel Stochastic Optimization with Wait-Avoiding Group Averaging
von: Li, Shigang, et al.
Veröffentlicht: (2020)
von: Li, Shigang, et al.
Veröffentlicht: (2020)
A Full Compression Pipeline for Green Federated Learning in Communication-Constrained Environments
von: Colybes, Elouan, et al.
Veröffentlicht: (2026)
von: Colybes, Elouan, et al.
Veröffentlicht: (2026)
Aergia: Leveraging Heterogeneity in Federated Learning Systems
von: Cox, Bart, et al.
Veröffentlicht: (2022)
von: Cox, Bart, et al.
Veröffentlicht: (2022)
Towards Optimal Heterogeneous Client Sampling in Multi-Model Federated Learning
von: Zhang, Haoran, et al.
Veröffentlicht: (2025)
von: Zhang, Haoran, et al.
Veröffentlicht: (2025)
Parameterizing Federated Continual Learning for Reproducible Research
von: Cox, Bart, et al.
Veröffentlicht: (2024)
von: Cox, Bart, et al.
Veröffentlicht: (2024)
Asynchronous Byzantine Federated Learning
von: Cox, Bart, et al.
Veröffentlicht: (2024)
von: Cox, Bart, et al.
Veröffentlicht: (2024)
Hyper-parameter Optimization for Federated Learning with Step-wise Adaptive Mechanism
von: Saadati, Yasaman, et al.
Veröffentlicht: (2024)
von: Saadati, Yasaman, et al.
Veröffentlicht: (2024)
Asynchronous Multi-Server Federated Learning for Geo-Distributed Clients
von: Zuo, Yuncong, et al.
Veröffentlicht: (2024)
von: Zuo, Yuncong, et al.
Veröffentlicht: (2024)
FLEdge: Benchmarking Federated Machine Learning Applications in Edge Computing Systems
von: Woisetschläger, Herbert, et al.
Veröffentlicht: (2023)
von: Woisetschläger, Herbert, et al.
Veröffentlicht: (2023)
WLB-LLM: Workload-Balanced 4D Parallelism for Large Language Model Training
von: Wang, Zheng, et al.
Veröffentlicht: (2025)
von: Wang, Zheng, et al.
Veröffentlicht: (2025)
Scheduling the Unschedulable: Taming Black-Box LLM Inference at Scale
von: Yuan, Renzhong, et al.
Veröffentlicht: (2026)
von: Yuan, Renzhong, et al.
Veröffentlicht: (2026)
FedPLT: Scalable, Resource-Efficient, and Heterogeneity-Aware Federated Learning via Partial Layer Training
von: Dabaja, Ahmad, et al.
Veröffentlicht: (2026)
von: Dabaja, Ahmad, et al.
Veröffentlicht: (2026)
DAGER: Exact Gradient Inversion for Large Language Models
von: Petrov, Ivo, et al.
Veröffentlicht: (2024)
von: Petrov, Ivo, et al.
Veröffentlicht: (2024)
Bridging Generalization Gap of Heterogeneous Federated Clients Using Generative Models
von: Niu, Ziru, et al.
Veröffentlicht: (2025)
von: Niu, Ziru, et al.
Veröffentlicht: (2025)
A Taxonomy and Resolution Strategy for Client-Level Disagreements in Federated Learning
von: Rosendal, Daan, et al.
Veröffentlicht: (2026)
von: Rosendal, Daan, et al.
Veröffentlicht: (2026)
From Detection to Recovery: Operational Analysis on LLM Pre-training with 504 GPUs
von: Kang, Daemyung, et al.
Veröffentlicht: (2026)
von: Kang, Daemyung, et al.
Veröffentlicht: (2026)
Federated Fine-Tuning of LLMs on the Very Edge: The Good, the Bad, the Ugly
von: Woisetschläger, Herbert, et al.
Veröffentlicht: (2023)
von: Woisetschläger, Herbert, et al.
Veröffentlicht: (2023)
Mobile Traffic Prediction at the Edge Through Distributed and Deep Transfer Learning
von: Petrella, Alfredo, et al.
Veröffentlicht: (2023)
von: Petrella, Alfredo, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Near-Optimal Sparse Allreduce for Distributed Deep Learning
von: Li, Shigang, et al.
Veröffentlicht: (2022) -
AutoDDL: Automatic Distributed Deep Learning with Near-Optimal Bandwidth Cost
von: Chen, Jinfan, et al.
Veröffentlicht: (2023) -
SparkAttention: High-Performance Multi-Head Attention for Large Models on Volta GPU Architecture
von: Xu, Youxuan, et al.
Veröffentlicht: (2025) -
FlashSparse: Minimizing Computation Redundancy for Fast Sparse Matrix Multiplications on Tensor Cores
von: Shi, Jinliang, et al.
Veröffentlicht: (2024) -
Libra: Unleashing GPU Heterogeneity for High-Performance Sparse Matrix Multiplication
von: Shi, Jinliang, et al.
Veröffentlicht: (2025)