ZenFlow: Enabling Stall-Free Offloading Training via Asynchronous Updates
Fuente:
arXiv
Saved in:
| Main Authors: | Lan, Tingfeng, Wu, Yusen, Ma, Bin, Su, Zhaoyuan, Yang, Rui, Bicer, Tekin, Tanaka, Masahiro, Ruwase, Olatunji, Li, Dong, Cheng, Yue |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SuperOffload: Unleashing the Power of Large-Scale LLM Training on Superchips
by: Lian, Xinyu, et al.
Published: (2025)
by: Lian, Xinyu, et al.
Published: (2025)
Flex-MIG: Enabling Distributed Execution on MIG
by: Kim, Myeongsu, et al.
Published: (2025)
by: Kim, Myeongsu, et al.
Published: (2025)
DPDPU: Data Processing with DPUs
by: Hu, Jiasheng, et al.
Published: (2024)
by: Hu, Jiasheng, et al.
Published: (2024)
DeepCompile: A Compiler-Driven Approach to Optimizing Distributed Deep Learning Training
by: Tanaka, Masahiro, et al.
Published: (2025)
by: Tanaka, Masahiro, et al.
Published: (2025)
push0: Scalable and Fault-Tolerant Orchestration for Zero-Knowledge Proof Generation
by: Ahmadvand, Mohsen, et al.
Published: (2026)
by: Ahmadvand, Mohsen, et al.
Published: (2026)
Serverless Cold Starts and Where to Find Them
by: Joosen, Artjom, et al.
Published: (2024)
by: Joosen, Artjom, et al.
Published: (2024)
nvidia-pcm: A D-Bus-Driven Platform Configuration Manager for OpenBMC Environments
by: Singh, Harinder
Published: (2026)
by: Singh, Harinder
Published: (2026)
AutoSP: Unlocking Long-Context LLM Training Via Compiler-Based Sequence Parallelism
by: Gupta, Ahan, et al.
Published: (2026)
by: Gupta, Ahan, et al.
Published: (2026)
GridPilot: Real-Time Grid-Responsive Control for AI Supercomputers
by: Constantinescu, Denisa-Andreea, et al.
Published: (2026)
by: Constantinescu, Denisa-Andreea, et al.
Published: (2026)
Alea-BFT: Practical Asynchronous Byzantine Fault Tolerance
by: Antunes, Diogo S., et al.
Published: (2024)
by: Antunes, Diogo S., et al.
Published: (2024)
Universal Checkpointing: A Flexible and Efficient Distributed Checkpointing System for Large-Scale DNN Training with Reconfigurable Parallelis
by: Lian, Xinyu, et al.
Published: (2024)
by: Lian, Xinyu, et al.
Published: (2024)
Training Ultra Long Context Language Model with Fully Pipelined Distributed Transformer
by: Yao, Jinghan, et al.
Published: (2024)
by: Yao, Jinghan, et al.
Published: (2024)
Efficiently Scheduling Parallel DAG Tasks on Identical Multiprocessors
by: Lendve, Shardul, et al.
Published: (2024)
by: Lendve, Shardul, et al.
Published: (2024)
TStore: Rethinking AI Model Hub with Tensor-Centric Compression
by: Lan, Tingfeng, et al.
Published: (2026)
by: Lan, Tingfeng, et al.
Published: (2026)
Service Discovery-Based Hybrid Network Middleware for Efficient Communication in Distributed Robotic Systems
by: Sang, Shiyao, et al.
Published: (2025)
by: Sang, Shiyao, et al.
Published: (2025)
Directives for Function Offloading in 5G Networks Based on a Performance Characteristics Analysis
by: Dettinger, Falk, et al.
Published: (2025)
by: Dettinger, Falk, et al.
Published: (2025)
FCDP: Fully Cached Data Parallel for Communication-Avoiding Large-Scale Training
by: Park, Gyeongseo, et al.
Published: (2026)
by: Park, Gyeongseo, et al.
Published: (2026)
Evaluating Large Language Models for Workload Mapping and Scheduling in Heterogeneous HPC Systems
by: Sharma, Aasish Kumar, et al.
Published: (2025)
by: Sharma, Aasish Kumar, et al.
Published: (2025)
CoCoDiff: Optimizing Collective Communications for Distributed Diffusion Transformer Inference Under Ulysses Sequence Parallelism
by: Ma, Bin, et al.
Published: (2026)
by: Ma, Bin, et al.
Published: (2026)
GPU-Augmented OLAP Execution Engine: GPU Offloading
by: Chang, Ilsun
Published: (2025)
by: Chang, Ilsun
Published: (2025)
ZipLLM: Efficient LLM Storage via Model-Aware Synergistic Data Deduplication and Compression
by: Wang, Zirui, et al.
Published: (2025)
by: Wang, Zirui, et al.
Published: (2025)
Cross-Layer Energy Analysis of Multimodal Training on Grace Hopper Superchips
by: Ahmed, Mahmoud, et al.
Published: (2026)
by: Ahmed, Mahmoud, et al.
Published: (2026)
Melding the Serverless Control Plane with the Conventional Cluster Manager for Speed and Resource Efficiency
by: Kondrashov, Leonid, et al.
Published: (2025)
by: Kondrashov, Leonid, et al.
Published: (2025)
NotebookOS: A Replicated Notebook Platform for Interactive Training with On-Demand GPUs
by: Carver, Benjamin, et al.
Published: (2025)
by: Carver, Benjamin, et al.
Published: (2025)
Scalable Engine and the Performance of Different LLM Models in a SLURM based HPC architecture
by: Luiz, Anderson de Lima, et al.
Published: (2025)
by: Luiz, Anderson de Lima, et al.
Published: (2025)
ConfigSpec: Profiling-Based Configuration Selection for Distributed Edge--Cloud Speculative LLM Serving
by: Li, Xiangchen, et al.
Published: (2026)
by: Li, Xiangchen, et al.
Published: (2026)
WISP: Waste- and Interference-Suppressed Distributed Speculative LLM Serving at the Edge via Dynamic Drafting and SLO-Aware Batching
by: Li, Xiangchen, et al.
Published: (2026)
by: Li, Xiangchen, et al.
Published: (2026)
Domino: Eliminating Communication in LLM Training via Generic Tensor Slicing and Overlapping
by: Wang, Guanhua, et al.
Published: (2024)
by: Wang, Guanhua, et al.
Published: (2024)
Send: Objects, History, and Transactions in a Single-Verb Kernel
by: Goes, Christopher
Published: (2026)
by: Goes, Christopher
Published: (2026)
mLR: Scalable Laminography Reconstruction based on Memoization
by: Ma, Bin, et al.
Published: (2025)
by: Ma, Bin, et al.
Published: (2025)
Sky$^ε$-Tree: Embracing the Batch Updates of B$^ε$-trees through Access Port Parallelism on Skyrmion Racetrack Memory
by: Tsai, Yu-Shiang, et al.
Published: (2024)
by: Tsai, Yu-Shiang, et al.
Published: (2024)
OPTIMUMP2P: Fast and Reliable Gossiping in P2P Networks
by: Nicolaou, Nicolas, et al.
Published: (2025)
by: Nicolaou, Nicolas, et al.
Published: (2025)
A Treasure Trove of Performance: Analyzing the IO500 Submission Data
by: Kunkel, Julian, et al.
Published: (2026)
by: Kunkel, Julian, et al.
Published: (2026)
Limitless FaaS: Overcoming serverless functions execution time limits with invoke driven architecture and memory checkpoints
by: Andraca, Rodrigo Landa, et al.
Published: (2024)
by: Andraca, Rodrigo Landa, et al.
Published: (2024)
Building the Palmetto API: Adding granular permissions and caching to the Slurm REST API without sacrificing compatibility
by: Godfrey, Ben, et al.
Published: (2026)
by: Godfrey, Ben, et al.
Published: (2026)
Efficient and Reuseable Cloud Configuration Search Using Discovery Spaces
by: Johnston, Michael, et al.
Published: (2025)
by: Johnston, Michael, et al.
Published: (2025)
Scalability Evaluation of HPC Multi-GPU Training for ECG-based LLMs
by: Mileski, Dimitar, et al.
Published: (2025)
by: Mileski, Dimitar, et al.
Published: (2025)
Relaxation for Efficient Asynchronous Queues
by: Baldwin, Samuel, et al.
Published: (2025)
by: Baldwin, Samuel, et al.
Published: (2025)
FastPersist: Accelerating Model Checkpointing in Deep Learning
by: Wang, Guanhua, et al.
Published: (2024)
by: Wang, Guanhua, et al.
Published: (2024)
DDS: DPU-optimized Disaggregated Storage [Extended Report]
by: Zhang, Qizhen, et al.
Published: (2024)
by: Zhang, Qizhen, et al.
Published: (2024)
Similar Items
-
SuperOffload: Unleashing the Power of Large-Scale LLM Training on Superchips
by: Lian, Xinyu, et al.
Published: (2025) -
Flex-MIG: Enabling Distributed Execution on MIG
by: Kim, Myeongsu, et al.
Published: (2025) -
DPDPU: Data Processing with DPUs
by: Hu, Jiasheng, et al.
Published: (2024) -
DeepCompile: A Compiler-Driven Approach to Optimizing Distributed Deep Learning Training
by: Tanaka, Masahiro, et al.
Published: (2025) -
push0: Scalable and Fault-Tolerant Orchestration for Zero-Knowledge Proof Generation
by: Ahmadvand, Mohsen, et al.
Published: (2026)