FastPersist: Accelerating Model Checkpointing in Deep Learning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Guanhua, Ruwase, Olatunji, Xie, Bing, He, Yuxiong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Domino: Eliminating Communication in LLM Training via Generic Tensor Slicing and Overlapping
von: Wang, Guanhua, et al.
Veröffentlicht: (2024)
von: Wang, Guanhua, et al.
Veröffentlicht: (2024)
AutoSP: Unlocking Long-Context LLM Training Via Compiler-Based Sequence Parallelism
von: Gupta, Ahan, et al.
Veröffentlicht: (2026)
von: Gupta, Ahan, et al.
Veröffentlicht: (2026)
Democratizing AI: A Comparative Study in Deep Learning Efficiency and Future Trends in Computational Processing
von: Amin, Lisan Al, et al.
Veröffentlicht: (2026)
von: Amin, Lisan Al, et al.
Veröffentlicht: (2026)
Low-Rank GEMM: Efficient Matrix Multiplication via Low-Rank Approximation with FP8 Acceleration
von: Metere, Alfredo
Veröffentlicht: (2025)
von: Metere, Alfredo
Veröffentlicht: (2025)
Training Ultra Long Context Language Model with Fully Pipelined Distributed Transformer
von: Yao, Jinghan, et al.
Veröffentlicht: (2024)
von: Yao, Jinghan, et al.
Veröffentlicht: (2024)
Lightweight Software Kernels and Hardware Extensions for Efficient Sparse Deep Neural Networks on Microcontrollers
von: Daghero, Francesco, et al.
Veröffentlicht: (2025)
von: Daghero, Francesco, et al.
Veröffentlicht: (2025)
GhostServe: A Lightweight Checkpointing System in the Shadow for Fault-Tolerant LLM Serving
von: Jayakody, Shakya, et al.
Veröffentlicht: (2026)
von: Jayakody, Shakya, et al.
Veröffentlicht: (2026)
Universal Checkpointing: A Flexible and Efficient Distributed Checkpointing System for Large-Scale DNN Training with Reconfigurable Parallelis
von: Lian, Xinyu, et al.
Veröffentlicht: (2024)
von: Lian, Xinyu, et al.
Veröffentlicht: (2024)
Reliable Microservice Tail Latency Prediction via Decoupled Dual-Stream Learning and Gradient Modulation
von: Qian, Wenzhuo, et al.
Veröffentlicht: (2025)
von: Qian, Wenzhuo, et al.
Veröffentlicht: (2025)
STAlloc: Enhancing Memory Efficiency in Large-Scale Model Training with Spatio-Temporal Planning
von: Huang, Zixiao, et al.
Veröffentlicht: (2025)
von: Huang, Zixiao, et al.
Veröffentlicht: (2025)
Prism: Unleashing GPU Sharing for Cost-Efficient Multi-LLM Serving
von: Yu, Shan, et al.
Veröffentlicht: (2025)
von: Yu, Shan, et al.
Veröffentlicht: (2025)
Sometimes Painful but Certainly Promising: Feasibility and Trade-offs of Language Model Inference at the Edge
von: Abstreiter, Maximilian, et al.
Veröffentlicht: (2025)
von: Abstreiter, Maximilian, et al.
Veröffentlicht: (2025)
QPART: Adaptive Model Quantization and Dynamic Workload Balancing for Accuracy-aware Edge Inference
von: Li, Xiangchen, et al.
Veröffentlicht: (2025)
von: Li, Xiangchen, et al.
Veröffentlicht: (2025)
Towards Efficient Generative Large Language Model Serving: A Survey from Algorithms to Systems
von: Miao, Xupeng, et al.
Veröffentlicht: (2023)
von: Miao, Xupeng, et al.
Veröffentlicht: (2023)
GPU Kernel Optimization Beyond Full Builds: An LLM Framework with Minimal Executable Programs
von: Chu, Ruifan, et al.
Veröffentlicht: (2025)
von: Chu, Ruifan, et al.
Veröffentlicht: (2025)
Prompt-Aware Scheduling for Low-Latency LLM Serving
von: Tao, Yiheng, et al.
Veröffentlicht: (2025)
von: Tao, Yiheng, et al.
Veröffentlicht: (2025)
ProTrain: Efficient LLM Training via Memory-Aware Techniques
von: Yang, Hanmei, et al.
Veröffentlicht: (2024)
von: Yang, Hanmei, et al.
Veröffentlicht: (2024)
Compiler-First State Space Duality and Portable $O(1)$ Autoregressive Caching for Inference
von: Santoni, Cosmo
Veröffentlicht: (2026)
von: Santoni, Cosmo
Veröffentlicht: (2026)
Multi-Dimensional Autoscaling of Stream Processing Services on Edge Devices
von: Sedlak, Boris, et al.
Veröffentlicht: (2025)
von: Sedlak, Boris, et al.
Veröffentlicht: (2025)
The Illusion of Power Capping in LLM Decode: A Phase-Aware Energy Characterisation Across Attention Architectures
von: Ma, Bole, et al.
Veröffentlicht: (2026)
von: Ma, Bole, et al.
Veröffentlicht: (2026)
Record-Remix-Replay: Hierarchical GPU Kernel Optimization using Evolutionary Search
von: Nichols, Daniel, et al.
Veröffentlicht: (2026)
von: Nichols, Daniel, et al.
Veröffentlicht: (2026)
Syno: Structured Synthesis for Neural Operators
von: Zhuo, Yongqi, et al.
Veröffentlicht: (2024)
von: Zhuo, Yongqi, et al.
Veröffentlicht: (2024)
TrainMover: An Interruption-Resilient Runtime for ML Training
von: Lao, ChonLam, et al.
Veröffentlicht: (2024)
von: Lao, ChonLam, et al.
Veröffentlicht: (2024)
Longer Attention Span: Increasing Transformer Context Length with Sparse Graph Processing Techniques
von: Tomczak, Nathaniel, et al.
Veröffentlicht: (2025)
von: Tomczak, Nathaniel, et al.
Veröffentlicht: (2025)
Mixture of Experts with Mixture of Precisions for Tuning Quality of Service
von: Imani, HamidReza, et al.
Veröffentlicht: (2024)
von: Imani, HamidReza, et al.
Veröffentlicht: (2024)
Training Time Prediction for Mixed Precision-based Distributed Training
von: Kang, Minchul, et al.
Veröffentlicht: (2026)
von: Kang, Minchul, et al.
Veröffentlicht: (2026)
A Comparative Study of OpenMP Scheduling Algorithm Selection Strategies
von: Korndörfer, Jonas H. Müller, et al.
Veröffentlicht: (2025)
von: Korndörfer, Jonas H. Müller, et al.
Veröffentlicht: (2025)
MQ-GNN: A Multi-Queue Pipelined Architecture for Scalable and Efficient GNN Training
von: Ullah, Irfan, et al.
Veröffentlicht: (2026)
von: Ullah, Irfan, et al.
Veröffentlicht: (2026)
Cost-Efficient LLM Training with Lifetime-Aware Tensor Offloading via GPUDirect Storage
von: Yuan, Ziqi, et al.
Veröffentlicht: (2025)
von: Yuan, Ziqi, et al.
Veröffentlicht: (2025)
A 4D Hybrid Algorithm to Scale Parallel Training to Thousands of GPUs
von: Singh, Siddharth, et al.
Veröffentlicht: (2023)
von: Singh, Siddharth, et al.
Veröffentlicht: (2023)
Optimizing the Deployment of Tiny Transformers on Low-Power MCUs
von: Jung, Victor J. B., et al.
Veröffentlicht: (2024)
von: Jung, Victor J. B., et al.
Veröffentlicht: (2024)
OSCAR: Offline Spectral Covariance-Aware Rotation for 2-bit KV Cache Quantization
von: Zhou, Zhongzhu, et al.
Veröffentlicht: (2026)
von: Zhou, Zhongzhu, et al.
Veröffentlicht: (2026)
SCAR: Scheduling Multi-Model AI Workloads on Heterogeneous Multi-Chiplet Module Accelerators
von: Odema, Mohanad, et al.
Veröffentlicht: (2024)
von: Odema, Mohanad, et al.
Veröffentlicht: (2024)
Green or Fast? Learning to Balance Cold Starts and Idle Carbon in Serverless Computing
von: Sun, Bowen, et al.
Veröffentlicht: (2026)
von: Sun, Bowen, et al.
Veröffentlicht: (2026)
Agentic Plan Caching: Test-Time Memory for Fast and Cost-Efficient LLM Agents
von: Zhang, Qizheng, et al.
Veröffentlicht: (2025)
von: Zhang, Qizheng, et al.
Veröffentlicht: (2025)
Binary Bleed: Fast Distributed and Parallel Method for Automatic Model Selection
von: Barron, Ryan, et al.
Veröffentlicht: (2024)
von: Barron, Ryan, et al.
Veröffentlicht: (2024)
EdgeProfiler: A Fast Profiling Framework for Lightweight LLMs on Edge Using Analytical Model
von: Pinnock, Alyssa, et al.
Veröffentlicht: (2025)
von: Pinnock, Alyssa, et al.
Veröffentlicht: (2025)
Performance and Power: Systematic Evaluation of AI Workloads on Accelerators with CARAML
von: John, Chelsea Maria, et al.
Veröffentlicht: (2024)
von: John, Chelsea Maria, et al.
Veröffentlicht: (2024)
SuperOffload: Unleashing the Power of Large-Scale LLM Training on Superchips
von: Lian, Xinyu, et al.
Veröffentlicht: (2025)
von: Lian, Xinyu, et al.
Veröffentlicht: (2025)
Deep Optimizer States: Towards Scalable Training of Transformer Models Using Interleaved Offloading
von: Maurya, Avinash, et al.
Veröffentlicht: (2024)
von: Maurya, Avinash, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Domino: Eliminating Communication in LLM Training via Generic Tensor Slicing and Overlapping
von: Wang, Guanhua, et al.
Veröffentlicht: (2024) -
AutoSP: Unlocking Long-Context LLM Training Via Compiler-Based Sequence Parallelism
von: Gupta, Ahan, et al.
Veröffentlicht: (2026) -
Democratizing AI: A Comparative Study in Deep Learning Efficiency and Future Trends in Computational Processing
von: Amin, Lisan Al, et al.
Veröffentlicht: (2026) -
Low-Rank GEMM: Efficient Matrix Multiplication via Low-Rank Approximation with FP8 Acceleration
von: Metere, Alfredo
Veröffentlicht: (2025) -
Training Ultra Long Context Language Model with Fully Pipelined Distributed Transformer
von: Yao, Jinghan, et al.
Veröffentlicht: (2024)