Making Wide Stripes Practical: Cascaded Parity LRCs for Efficient Repair and High Reliability
Fuente:
arXiv
Guardado en:
| Autores principales: | Yu, Fan, Li, Guodong, Wu, Si, Fang, Weijun, Hu, Sihuang |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Cascadia: An Efficient Cascade Serving System for Large Language Models
por: Jiang, Youhe, et al.
Publicado: (2025)
por: Jiang, Youhe, et al.
Publicado: (2025)
Lumiere: Making Optimal BFT for Partial Synchrony Practical
por: Lewis-Pye, Andrew, et al.
Publicado: (2023)
por: Lewis-Pye, Andrew, et al.
Publicado: (2023)
CacheFlow: Efficient LLM Serving with 3D-Parallel KV Cache Restoration
por: Nian, Sean, et al.
Publicado: (2026)
por: Nian, Sean, et al.
Publicado: (2026)
Amortized Asynchronous Byzantine Reliable Broadcast with Optimal Resilience
por: Hu, Michael Yiqing, et al.
Publicado: (2026)
por: Hu, Michael Yiqing, et al.
Publicado: (2026)
An Efficient, Reliable and Observable Collective Communication Library in Large-scale GPU Training Clusters
por: Zhang, Mingjun, et al.
Publicado: (2025)
por: Zhang, Mingjun, et al.
Publicado: (2025)
Efficient Profit Maximization in Reliability Concerned Static Vehicular Cloud System
por: Sarkar, Suvarthi, et al.
Publicado: (2023)
por: Sarkar, Suvarthi, et al.
Publicado: (2023)
To Repair or Not to Repair: Assessing Fault Resilience in MPI Stencil Applications
por: Rocco, Roberto, et al.
Publicado: (2024)
por: Rocco, Roberto, et al.
Publicado: (2024)
Training Overhead Ratio: A Practical Reliability Metric for Large Language Model Training Systems
por: Lu, Ning, et al.
Publicado: (2024)
por: Lu, Ning, et al.
Publicado: (2024)
New Wide Locally Recoverable Codes with Unified Locality
por: Xu, Liangliang, et al.
Publicado: (2025)
por: Xu, Liangliang, et al.
Publicado: (2025)
IsoSched: Preemptive Tile Cascaded Scheduling of Multi-DNN via Subgraph Isomorphism
por: Zhao, Boran, et al.
Publicado: (2025)
por: Zhao, Boran, et al.
Publicado: (2025)
Surviving Partial Rank Failures in Wide Expert-Parallel MoE Inference
por: Sun, Xun, et al.
Publicado: (2026)
por: Sun, Xun, et al.
Publicado: (2026)
Making Serverless Computing Extensible: A Case Study of Serverless Data Analytics
por: Yu, Minchen, et al.
Publicado: (2025)
por: Yu, Minchen, et al.
Publicado: (2025)
CascadeInfer: Length-Aware Scheduling of LLM Serving with Low Latency and Load Balancing
por: Yuan, Yitao, et al.
Publicado: (2025)
por: Yuan, Yitao, et al.
Publicado: (2025)
Optimistic, Signature-Free Reliable Broadcast and Its Applications
por: Shrestha, Nibesh, et al.
Publicado: (2025)
por: Shrestha, Nibesh, et al.
Publicado: (2025)
SpotVista: Availability-Aware Recommendation System for Reliable and Cost-Efficient Multi-Node Spot Instances
por: Kim, Taeyoon, et al.
Publicado: (2026)
por: Kim, Taeyoon, et al.
Publicado: (2026)
Parallel Collaborative ADMM Privacy Computing and Adaptive GPU Acceleration for Distributed Edge Networks
por: Xia, Mengchun, et al.
Publicado: (2026)
por: Xia, Mengchun, et al.
Publicado: (2026)
ElasticMoE: An Efficient Auto Scaling Method for Mixture-of-Experts Models
por: Singh, Gursimran, et al.
Publicado: (2025)
por: Singh, Gursimran, et al.
Publicado: (2025)
Automated, Reliable, and Efficient Continental-Scale Replication of 7.3 Petabytes of Climate Simulation Data: A Case Study
por: Lacinski, Lukasz, et al.
Publicado: (2024)
por: Lacinski, Lukasz, et al.
Publicado: (2024)
Fault-Tolerant Hybrid-Parallel Training at Scale with Reliable and Efficient In-memory Checkpointing
por: Wang, Yuxin, et al.
Publicado: (2023)
por: Wang, Yuxin, et al.
Publicado: (2023)
Dynamic Probabilistic Reliable Broadcast
por: Anikina, Veronika, et al.
Publicado: (2023)
por: Anikina, Veronika, et al.
Publicado: (2023)
Highly-Efficient Persistent FIFO Queues
por: Fatourou, Panagiota, et al.
Publicado: (2024)
por: Fatourou, Panagiota, et al.
Publicado: (2024)
Exact, Efficient, and Reliable Multi-Objective and Multi-Constrained IoT Workflow Scheduling in Edge-Hub-Cloud Cyber-Physical Systems
por: Kouloumpris, Andreas, et al.
Publicado: (2026)
por: Kouloumpris, Andreas, et al.
Publicado: (2026)
Pier: Efficient Large Language Model pretraining with Relaxed Global Communication
por: Fan, Shuyuan, et al.
Publicado: (2025)
por: Fan, Shuyuan, et al.
Publicado: (2025)
PALM: A Efficient Performance Simulator for Tiled Accelerators with Large-scale Model Training
por: Fang, Jiahao, et al.
Publicado: (2024)
por: Fang, Jiahao, et al.
Publicado: (2024)
OmniInfer: System-Wide Acceleration Techniques for Optimizing LLM Serving Throughput and Latency
por: Wang, Jun, et al.
Publicado: (2025)
por: Wang, Jun, et al.
Publicado: (2025)
Reliable Replication Protocols on SmartNICs
por: Katebzadeh, M. R. Siavash, et al.
Publicado: (2025)
por: Katebzadeh, M. R. Siavash, et al.
Publicado: (2025)
Joint$λ$: Orchestrating Serverless Workflows on Jointcloud FaaS Systems
por: Li, Rui, et al.
Publicado: (2025)
por: Li, Rui, et al.
Publicado: (2025)
M$^2$-MFP: A Multi-Scale and Multi-Level Memory Failure Prediction Framework for Reliable Cloud Infrastructure
por: Xie, Hongyi, et al.
Publicado: (2025)
por: Xie, Hongyi, et al.
Publicado: (2025)
GoodServe: Towards High-Goodput Serving of Agentic LLM Inferences over Heterogeneous Resources
por: Du, Boxiao, et al.
Publicado: (2026)
por: Du, Boxiao, et al.
Publicado: (2026)
FWeb3: A Practical Incentive-Aware Federated Learning Framework
por: Yan, Peishen, et al.
Publicado: (2026)
por: Yan, Peishen, et al.
Publicado: (2026)
HAP: Hybrid Adaptive Parallelism for Efficient Mixture-of-Experts Inference
por: Lin, Haoran, et al.
Publicado: (2025)
por: Lin, Haoran, et al.
Publicado: (2025)
ExpertWeave: Efficiently Serving Expert-Specialized Fine-Tuned Adapters at Scale
por: Shi, Ge, et al.
Publicado: (2025)
por: Shi, Ge, et al.
Publicado: (2025)
STRATUS: A Multi-agent System for Autonomous Reliability Engineering of Modern Clouds
por: Chen, Yinfang, et al.
Publicado: (2025)
por: Chen, Yinfang, et al.
Publicado: (2025)
Memory-Efficient Federated Fine-Tuning of Large Language Models via Layer Pruning
por: Wu, Yebo, et al.
Publicado: (2025)
por: Wu, Yebo, et al.
Publicado: (2025)
Distributed Consensus Network: A Modularized Communication Framework and Reliability Probabilistic Analysis
por: Li, Yuetai, et al.
Publicado: (2025)
por: Li, Yuetai, et al.
Publicado: (2025)
A Unified, Practical, and Understandable Model of Non-transactional Consistency Levels in Distributed Replication
por: Hu, Guanzhou, et al.
Publicado: (2024)
por: Hu, Guanzhou, et al.
Publicado: (2024)
Cicada: A Pipeline-Efficient Approach to Serverless Inference with Decoupled Management
por: Wu, Z., et al.
Publicado: (2025)
por: Wu, Z., et al.
Publicado: (2025)
CascadeServe: Unlocking Model Cascades for Inference Serving
por: Kossmann, Ferdi, et al.
Publicado: (2024)
por: Kossmann, Ferdi, et al.
Publicado: (2024)
Reliable Communication in Hybrid Authentication and Trust Models
por: Chotkan, Rowdy, et al.
Publicado: (2024)
por: Chotkan, Rowdy, et al.
Publicado: (2024)
S-HPLB: Efficient LLM Attention Serving via Sparsity-Aware Head Parallelism Load Balance
por: Liu, Di, et al.
Publicado: (2026)
por: Liu, Di, et al.
Publicado: (2026)
Ejemplares similares
-
Cascadia: An Efficient Cascade Serving System for Large Language Models
por: Jiang, Youhe, et al.
Publicado: (2025) -
Lumiere: Making Optimal BFT for Partial Synchrony Practical
por: Lewis-Pye, Andrew, et al.
Publicado: (2023) -
CacheFlow: Efficient LLM Serving with 3D-Parallel KV Cache Restoration
por: Nian, Sean, et al.
Publicado: (2026) -
Amortized Asynchronous Byzantine Reliable Broadcast with Optimal Resilience
por: Hu, Michael Yiqing, et al.
Publicado: (2026) -
An Efficient, Reliable and Observable Collective Communication Library in Large-scale GPU Training Clusters
por: Zhang, Mingjun, et al.
Publicado: (2025)