GhostServe: A Lightweight Checkpointing System in the Shadow for Fault-Tolerant LLM Serving
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Jayakody, Shakya, Zhao, Youpeng, Nehate, Chinmay Dhanraj, Wang, Jun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
EdgeProfiler: A Fast Profiling Framework for Lightweight LLMs on Edge Using Analytical Model
von: Pinnock, Alyssa, et al.
Veröffentlicht: (2025)
von: Pinnock, Alyssa, et al.
Veröffentlicht: (2025)
Prompt-Aware Scheduling for Low-Latency LLM Serving
von: Tao, Yiheng, et al.
Veröffentlicht: (2025)
von: Tao, Yiheng, et al.
Veröffentlicht: (2025)
CoServe: Efficient Collaboration-of-Experts (CoE) Model Inference with Limited Memory
von: Suo, Jiashun, et al.
Veröffentlicht: (2025)
von: Suo, Jiashun, et al.
Veröffentlicht: (2025)
Prism: Unleashing GPU Sharing for Cost-Efficient Multi-LLM Serving
von: Yu, Shan, et al.
Veröffentlicht: (2025)
von: Yu, Shan, et al.
Veröffentlicht: (2025)
Fault-Tolerant Hybrid-Parallel Training at Scale with Reliable and Efficient In-memory Checkpointing
von: Wang, Yuxin, et al.
Veröffentlicht: (2023)
von: Wang, Yuxin, et al.
Veröffentlicht: (2023)
BurstGPT: A Real-world Workload Dataset to Optimize LLM Serving Systems
von: Wang, Yuxin, et al.
Veröffentlicht: (2024)
von: Wang, Yuxin, et al.
Veröffentlicht: (2024)
Towards Efficient Generative Large Language Model Serving: A Survey from Algorithms to Systems
von: Miao, Xupeng, et al.
Veröffentlicht: (2023)
von: Miao, Xupeng, et al.
Veröffentlicht: (2023)
Serving Chain-structured Jobs with Large Memory Footprints with Application to Large Foundation Model Serving
von: Sun, Tingyang, et al.
Veröffentlicht: (2026)
von: Sun, Tingyang, et al.
Veröffentlicht: (2026)
Revisiting Disaggregated Large Language Model Serving for Performance and Energy Implications
von: Li, Jiaxi, et al.
Veröffentlicht: (2025)
von: Li, Jiaxi, et al.
Veröffentlicht: (2025)
LServe: Efficient Long-sequence LLM Serving with Unified Sparse Attention
von: Yang, Shang, et al.
Veröffentlicht: (2025)
von: Yang, Shang, et al.
Veröffentlicht: (2025)
Stabl: Blockchain Fault Tolerance
von: Gramoli, Vincent, et al.
Veröffentlicht: (2024)
von: Gramoli, Vincent, et al.
Veröffentlicht: (2024)
VibeServe: Can AI Agents Build Bespoke LLM Serving Systems?
von: Kamahori, Keisuke, et al.
Veröffentlicht: (2026)
von: Kamahori, Keisuke, et al.
Veröffentlicht: (2026)
FastPersist: Accelerating Model Checkpointing in Deep Learning
von: Wang, Guanhua, et al.
Veröffentlicht: (2024)
von: Wang, Guanhua, et al.
Veröffentlicht: (2024)
BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures
von: Hu, Xiannan, et al.
Veröffentlicht: (2025)
von: Hu, Xiannan, et al.
Veröffentlicht: (2025)
TinyServe: Query-Aware Cache Selection for Efficient LLM Serving
von: Liu, Dong, et al.
Veröffentlicht: (2025)
von: Liu, Dong, et al.
Veröffentlicht: (2025)
TierCheck: Tiered Checkpointing for Fault Tolerance in Large Language Model Training
von: Han, Shujie, et al.
Veröffentlicht: (2026)
von: Han, Shujie, et al.
Veröffentlicht: (2026)
StreamServe: Adaptive Speculative Flows for Low-Latency Disaggregated LLM Serving
von: Kumar, Satyam, et al.
Veröffentlicht: (2026)
von: Kumar, Satyam, et al.
Veröffentlicht: (2026)
EPD-Serve: A Flexible Multimodal EPD Disaggregation Inference Serving System On Ascend
von: Bai, Fan, et al.
Veröffentlicht: (2026)
von: Bai, Fan, et al.
Veröffentlicht: (2026)
BucketServe: Bucket-Based Dynamic Batching for Smart and Efficient LLM Inference Serving
von: Zheng, Wanyi, et al.
Veröffentlicht: (2025)
von: Zheng, Wanyi, et al.
Veröffentlicht: (2025)
LAPS: A Length-Aware-Prefill LLM Serving System
von: She, Jianshu, et al.
Veröffentlicht: (2026)
von: She, Jianshu, et al.
Veröffentlicht: (2026)
DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference
von: Jeong, Bodon, et al.
Veröffentlicht: (2026)
von: Jeong, Bodon, et al.
Veröffentlicht: (2026)
KunServe: Parameter-centric Memory Management for Efficient Memory Overloading Handling in LLM Serving
von: Cheng, Rongxin, et al.
Veröffentlicht: (2024)
von: Cheng, Rongxin, et al.
Veröffentlicht: (2024)
DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers
von: Maurya, Avinash, et al.
Veröffentlicht: (2026)
von: Maurya, Avinash, et al.
Veröffentlicht: (2026)
FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines
von: He, Jiaao, et al.
Veröffentlicht: (2024)
von: He, Jiaao, et al.
Veröffentlicht: (2024)
Shared Memory-contention-aware Concurrent DNN Execution for Diversely Heterogeneous System-on-Chips
von: Dagli, Ismet, et al.
Veröffentlicht: (2023)
von: Dagli, Ismet, et al.
Veröffentlicht: (2023)
LLM Inference Serving: Survey of Recent Advances and Opportunities
von: Li, Baolin, et al.
Veröffentlicht: (2024)
von: Li, Baolin, et al.
Veröffentlicht: (2024)
ModServe: Modality- and Stage-Aware Resource Disaggregation for Scalable Multimodal Model Serving
von: Qiu, Haoran, et al.
Veröffentlicht: (2025)
von: Qiu, Haoran, et al.
Veröffentlicht: (2025)
SkyServe: Serving AI Models across Regions and Clouds with Spot Instances
von: Mao, Ziming, et al.
Veröffentlicht: (2024)
von: Mao, Ziming, et al.
Veröffentlicht: (2024)
Scepsy: Serving Agentic Workflows Using Aggregate LLM Pipelines
von: Wagenländer, Marcel, et al.
Veröffentlicht: (2026)
von: Wagenländer, Marcel, et al.
Veröffentlicht: (2026)
PALS: Power-Aware LLM Serving for Mixture-of-Experts Models
von: Hankendi, Can, et al.
Veröffentlicht: (2026)
von: Hankendi, Can, et al.
Veröffentlicht: (2026)
ParaLog: Consistent Host-side Logging for Parallel Checkpoints
von: Chien, Steven W. D., et al.
Veröffentlicht: (2024)
von: Chien, Steven W. D., et al.
Veröffentlicht: (2024)
On Evaluating Performance of LLM Inference Serving Systems
von: Agrawal, Amey, et al.
Veröffentlicht: (2025)
von: Agrawal, Amey, et al.
Veröffentlicht: (2025)
PlanetServe: A Decentralized, Scalable, and Privacy-Preserving Overlay for Democratizing Large Language Model Serving
von: Fang, Fei, et al.
Veröffentlicht: (2025)
von: Fang, Fei, et al.
Veröffentlicht: (2025)
Lightweight Software Kernels and Hardware Extensions for Efficient Sparse Deep Neural Networks on Microcontrollers
von: Daghero, Francesco, et al.
Veröffentlicht: (2025)
von: Daghero, Francesco, et al.
Veröffentlicht: (2025)
Can Large Language Models Predict Parallel Code Performance?
von: Bolet, Gregory, et al.
Veröffentlicht: (2025)
von: Bolet, Gregory, et al.
Veröffentlicht: (2025)
oneDAL Optimization for ARM Scalable Vector Extension: Maximizing Efficiency for High-Performance Data Science
von: Sharma, Chandan, et al.
Veröffentlicht: (2025)
von: Sharma, Chandan, et al.
Veröffentlicht: (2025)
AGOCS -- Accurate Google Cloud Simulator Framework
von: Sliwko, Leszek, et al.
Veröffentlicht: (2025)
von: Sliwko, Leszek, et al.
Veröffentlicht: (2025)
Standardized Methods and Recommendations for Green Federated Learning
von: Tapp, Austin, et al.
Veröffentlicht: (2026)
von: Tapp, Austin, et al.
Veröffentlicht: (2026)
Confidential Computing on NVIDIA Hopper GPUs: A Performance Benchmark Study
von: Zhu, Jianwei, et al.
Veröffentlicht: (2024)
von: Zhu, Jianwei, et al.
Veröffentlicht: (2024)
Binary Bleed: Fast Distributed and Parallel Method for Automatic Model Selection
von: Barron, Ryan, et al.
Veröffentlicht: (2024)
von: Barron, Ryan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
EdgeProfiler: A Fast Profiling Framework for Lightweight LLMs on Edge Using Analytical Model
von: Pinnock, Alyssa, et al.
Veröffentlicht: (2025) -
Prompt-Aware Scheduling for Low-Latency LLM Serving
von: Tao, Yiheng, et al.
Veröffentlicht: (2025) -
CoServe: Efficient Collaboration-of-Experts (CoE) Model Inference with Limited Memory
von: Suo, Jiashun, et al.
Veröffentlicht: (2025) -
Prism: Unleashing GPU Sharing for Cost-Efficient Multi-LLM Serving
von: Yu, Shan, et al.
Veröffentlicht: (2025) -
Fault-Tolerant Hybrid-Parallel Training at Scale with Reliable and Efficient In-memory Checkpointing
von: Wang, Yuxin, et al.
Veröffentlicht: (2023)