FailSafe: High-performance Resilient Serving
Fuente:
arXiv
Saved in:
| Main Authors: | Xu, Ziyi, Xie, Zhiqiang, Gandhi, Swapnil, Kozyrakis, Christos |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Sparse Checkpointing for Fast and Reliable MoE Training
by: Gandhi, Swapnil, et al.
Published: (2024)
by: Gandhi, Swapnil, et al.
Published: (2024)
ReCycle: Resilient Training of Large DNNs using Pipeline Adaptation
by: Gandhi, Swapnil, et al.
Published: (2024)
by: Gandhi, Swapnil, et al.
Published: (2024)
Regulating Branch Parallelism in LLM Serving
by: Gandhi, Swapnil, et al.
Published: (2026)
by: Gandhi, Swapnil, et al.
Published: (2026)
Strata: Hierarchical Context Caching for Long Context Language Model Serving
by: Xie, Zhiqiang, et al.
Published: (2025)
by: Xie, Zhiqiang, et al.
Published: (2025)
SYMI: Efficient Mixture-of-Experts Training via Model and Optimizer State Decoupling
by: Skiadopoulos, Athinagoras, et al.
Published: (2025)
by: Skiadopoulos, Athinagoras, et al.
Published: (2025)
FailLite: Failure-Resilient Model Serving for Resource-Constrained Edge Environments
by: Wu, Li, et al.
Published: (2025)
by: Wu, Li, et al.
Published: (2025)
ThunderServe: High-performance and Cost-efficient LLM Serving in Cloud Environments
by: Jiang, Youhe, et al.
Published: (2025)
by: Jiang, Youhe, et al.
Published: (2025)
Fail-Closed Lowering of Resident KV Claims onto LLM Serving Runtimes
by: Stepanek, Lukas
Published: (2026)
by: Stepanek, Lukas
Published: (2026)
cedar: Optimized and Unified Machine Learning Input Data Pipelines
by: Zhao, Mark, et al.
Published: (2024)
by: Zhao, Mark, et al.
Published: (2024)
Cloud Atlas: Efficient Fault Localization for Cloud Systems using Language Models and Causal Insight
by: Xie, Zhiqiang, et al.
Published: (2024)
by: Xie, Zhiqiang, et al.
Published: (2024)
MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool
by: Hu, Cunchen, et al.
Published: (2024)
by: Hu, Cunchen, et al.
Published: (2024)
DeepServe: Serverless Large Language Model Serving at Scale
by: Hu, Junhao, et al.
Published: (2025)
by: Hu, Junhao, et al.
Published: (2025)
GoodServe: Towards High-Goodput Serving of Agentic LLM Inferences over Heterogeneous Resources
by: Du, Boxiao, et al.
Published: (2026)
by: Du, Boxiao, et al.
Published: (2026)
BrownoutServe: SLO-Aware Inference Serving under Bursty Workloads for MoE-based LLMs
by: Hu, Jianmin, et al.
Published: (2025)
by: Hu, Jianmin, et al.
Published: (2025)
TENT: A Declarative Slice Spraying Engine for Performant and Resilient Data Movement in Disaggregated LLM Serving
by: Ren, Feng, et al.
Published: (2026)
by: Ren, Feng, et al.
Published: (2026)
AlignedServe: Orchestrating Prefix-aware Batching to Build a High-throughput and Computing-efficient LLM Serving System
by: Bai, Fengyao, et al.
Published: (2026)
by: Bai, Fengyao, et al.
Published: (2026)
AI Metropolis: Scaling Large Language Model-based Multi-Agent Simulation with Out-of-order Execution
by: Xie, Zhiqiang, et al.
Published: (2024)
by: Xie, Zhiqiang, et al.
Published: (2024)
Cloud Native System for LLM Inference Serving
by: Xu, Minxian, et al.
Published: (2025)
by: Xu, Minxian, et al.
Published: (2025)
BanaServe: Unified KV Cache and Dynamic Module Migration for Balancing Disaggregated LLM Serving in AI Infrastructure
by: He, Yiyuan, et al.
Published: (2025)
by: He, Yiyuan, et al.
Published: (2025)
EdgeServing: Deadline-Aware Multi-DNN Serving at the Edge
by: Cao, Jiahe, et al.
Published: (2026)
by: Cao, Jiahe, et al.
Published: (2026)
Slice-Level Scheduling for High Throughput and Load Balanced LLM Serving
by: Cheng, Ke, et al.
Published: (2024)
by: Cheng, Ke, et al.
Published: (2024)
UELLM: A Unified and Efficient Approach for LLM Inference Serving
by: He, Yiyuan, et al.
Published: (2024)
by: He, Yiyuan, et al.
Published: (2024)
DynaServe: Unified and Elastic Execution for Dynamic Disaggregated LLM Serving
by: Ruan, Chaoyi, et al.
Published: (2025)
by: Ruan, Chaoyi, et al.
Published: (2025)
TridentServe: A Stage-level Serving System for Diffusion Pipelines
by: Xia, Yifei, et al.
Published: (2025)
by: Xia, Yifei, et al.
Published: (2025)
MuxServe: Flexible Spatial-Temporal Multiplexing for Multiple LLM Serving
by: Duan, Jiangfei, et al.
Published: (2024)
by: Duan, Jiangfei, et al.
Published: (2024)
ServeGen: Workload Characterization and Generation of Large Language Model Serving in Production
by: Xiang, Yuxing, et al.
Published: (2025)
by: Xiang, Yuxing, et al.
Published: (2025)
EconoServe: Maximizing Multi-Resource Utilization with SLO Guarantees in LLM Serving
by: Shen, Haiying, et al.
Published: (2024)
by: Shen, Haiying, et al.
Published: (2024)
Hetis: Serving LLMs in Heterogeneous GPU Clusters with Fine-grained and Dynamic Parallelism
by: Mo, Zizhao, et al.
Published: (2025)
by: Mo, Zizhao, et al.
Published: (2025)
PolyServe: Efficient Multi-SLO Serving at Scale
by: Zhu, Kan, et al.
Published: (2025)
by: Zhu, Kan, et al.
Published: (2025)
KnapsackLB: Enabling Performance-Aware Layer-4 Load Balancing
by: Gandhi, Rohan, et al.
Published: (2024)
by: Gandhi, Rohan, et al.
Published: (2024)
Serving Hybrid LLM Loads with SLO Guarantees Using CPU-GPU Attention Piggybacking
by: Mo, Zizhao, et al.
Published: (2026)
by: Mo, Zizhao, et al.
Published: (2026)
NanoFlow: Towards Optimal Large Language Model Serving Throughput
by: Zhu, Kan, et al.
Published: (2024)
by: Zhu, Kan, et al.
Published: (2024)
HydraServe: Minimizing Cold Start Latency for Serverless LLM Serving in Public Clouds
by: Lou, Chiheng, et al.
Published: (2025)
by: Lou, Chiheng, et al.
Published: (2025)
SparseServe: Unlocking Parallelism for Dynamic Sparse Attention in Long-Context LLM Serving
by: Zhou, Qihui, et al.
Published: (2025)
by: Zhou, Qihui, et al.
Published: (2025)
DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving
by: Zhong, Yinmin, et al.
Published: (2024)
by: Zhong, Yinmin, et al.
Published: (2024)
Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling
by: Wu, Jingfeng, et al.
Published: (2025)
by: Wu, Jingfeng, et al.
Published: (2025)
SageServe: Optimizing LLM Serving on Cloud Data Centers with Forecast Aware Auto-Scaling
by: Jaiswal, Shashwat, et al.
Published: (2025)
by: Jaiswal, Shashwat, et al.
Published: (2025)
EcoServe: Enabling Cost-effective LLM Serving with Proactive Intra- and Inter-Instance Orchestration
by: Du, Jiangsu, et al.
Published: (2025)
by: Du, Jiangsu, et al.
Published: (2025)
CaraServe: CPU-Assisted and Rank-Aware LoRA Serving for Generative LLM Inference
by: Li, Suyi, et al.
Published: (2024)
by: Li, Suyi, et al.
Published: (2024)
DiffServe: Efficiently Serving Text-to-Image Diffusion Models with Query-Aware Model Scaling
by: Ahmad, Sohaib, et al.
Published: (2024)
by: Ahmad, Sohaib, et al.
Published: (2024)
Similar Items
-
Sparse Checkpointing for Fast and Reliable MoE Training
by: Gandhi, Swapnil, et al.
Published: (2024) -
ReCycle: Resilient Training of Large DNNs using Pipeline Adaptation
by: Gandhi, Swapnil, et al.
Published: (2024) -
Regulating Branch Parallelism in LLM Serving
by: Gandhi, Swapnil, et al.
Published: (2026) -
Strata: Hierarchical Context Caching for Long Context Language Model Serving
by: Xie, Zhiqiang, et al.
Published: (2025) -
SYMI: Efficient Mixture-of-Experts Training via Model and Optimizer State Decoupling
by: Skiadopoulos, Athinagoras, et al.
Published: (2025)