CoCoI: Distributed Coded Inference System for Straggler Mitigation
Fuente:
arXiv
Salvato in:
| Autori principali: | Liu, Xing, Huang, Chao, Tang, Ming |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Exploiting Stragglers in Distributed Computing Systems with Task Grouping
di: Adikari, Tharindu, et al.
Pubblicazione: (2024)
di: Adikari, Tharindu, et al.
Pubblicazione: (2024)
FlowSpec: Continuous Pipelined Speculative Decoding for Efficient Distributed LLM Inference
di: Liu, Xing, et al.
Pubblicazione: (2025)
di: Liu, Xing, et al.
Pubblicazione: (2025)
dSTAR: Straggler Tolerant and Byzantine Resilient Distributed SGD
di: Yan, Jiahe, et al.
Pubblicazione: (2024)
di: Yan, Jiahe, et al.
Pubblicazione: (2024)
Sparsity-Preserving Encodings for Straggler-Optimal Distributed Matrix Computations at the Edge
di: Das, Anindya Bijoy, et al.
Pubblicazione: (2024)
di: Das, Anindya Bijoy, et al.
Pubblicazione: (2024)
Distributed Learning based on 1-Bit Gradient Coding in the Presence of Stragglers
di: Li, Chengxi, et al.
Pubblicazione: (2024)
di: Li, Chengxi, et al.
Pubblicazione: (2024)
CoCoDiff: Optimizing Collective Communications for Distributed Diffusion Transformer Inference Under Ulysses Sequence Parallelism
di: Ma, Bin, et al.
Pubblicazione: (2026)
di: Ma, Bin, et al.
Pubblicazione: (2026)
Optimal Server Selection for Straggler Mitigation
di: Badita, Ajay, et al.
Pubblicazione: (2019)
di: Badita, Ajay, et al.
Pubblicazione: (2019)
Straggler Tolerant and Resilient DL Training on Homogeneous GPUs
di: Zhang, Zeyu, et al.
Pubblicazione: (2025)
di: Zhang, Zeyu, et al.
Pubblicazione: (2025)
General Coded Computing in a Probabilistic Straggler Regime
di: Moradi, Parsa, et al.
Pubblicazione: (2025)
di: Moradi, Parsa, et al.
Pubblicazione: (2025)
Adaptra: Straggler-Resilient Hybrid-Parallel Training with Pipeline Adaptation
di: Wu, Tianyuan, et al.
Pubblicazione: (2025)
di: Wu, Tianyuan, et al.
Pubblicazione: (2025)
FALCON: Pinpointing and Mitigating Stragglers for Large-Scale Hybrid-Parallel Training
di: Wu, Tianyuan, et al.
Pubblicazione: (2024)
di: Wu, Tianyuan, et al.
Pubblicazione: (2024)
Uncoded Download in Lagrange-Coded Elastic Computing with Straggler Tolerance
di: Zhong, Xi, et al.
Pubblicazione: (2025)
di: Zhong, Xi, et al.
Pubblicazione: (2025)
LLM-CoOpt: A Co-Design and Optimization Framework for Efficient LLM Inference on Heterogeneous Platforms
di: Kong, Jie, et al.
Pubblicazione: (2026)
di: Kong, Jie, et al.
Pubblicazione: (2026)
ACE-GNN: Adaptive GNN Co-Inference with System-Aware Scheduling in Dynamic Edge Environments
di: Zhou, Ao, et al.
Pubblicazione: (2025)
di: Zhou, Ao, et al.
Pubblicazione: (2025)
CoGenT: A Content-oriented Generative-hit Framework for Content Delivery Networks
di: Wang, Peng, et al.
Pubblicazione: (2024)
di: Wang, Peng, et al.
Pubblicazione: (2024)
COSMIC: Enabling Full-Stack Co-Design and Optimization of Distributed Machine Learning Systems
di: Raju, Aditi, et al.
Pubblicazione: (2025)
di: Raju, Aditi, et al.
Pubblicazione: (2025)
Joint Optimization of Offloading, Batching and DVFS for Multiuser Co-Inference
di: Xu, Yaodan, et al.
Pubblicazione: (2025)
di: Xu, Yaodan, et al.
Pubblicazione: (2025)
Hiding Communication Cost in Distributed LLM Training via Micro-batch Co-execution
di: Wang, Haiquan, et al.
Pubblicazione: (2024)
di: Wang, Haiquan, et al.
Pubblicazione: (2024)
Biased Compression in Gradient Coding for Distributed Learning
di: Li, Chengxi, et al.
Pubblicazione: (2026)
di: Li, Chengxi, et al.
Pubblicazione: (2026)
Efficient AllReduce with Stragglers
di: Devraj, Arjun, et al.
Pubblicazione: (2025)
di: Devraj, Arjun, et al.
Pubblicazione: (2025)
Cooperative Gradient Coding
di: Weng, Shudi, et al.
Pubblicazione: (2025)
di: Weng, Shudi, et al.
Pubblicazione: (2025)
PROBE: Co-Balancing Computation and Communication in MoE Inference via Real-Time Predictive Prefetching
di: Zhu, Qianchao, et al.
Pubblicazione: (2026)
di: Zhu, Qianchao, et al.
Pubblicazione: (2026)
Sutradhara: An Intelligent Orchestrator-Engine Co-design for Tool-based Agentic Inference
di: Biswas, Anish, et al.
Pubblicazione: (2026)
di: Biswas, Anish, et al.
Pubblicazione: (2026)
OnePiece: A Large-Scale Distributed Inference System with RDMA for Complex AI-Generated Content (AIGC) Workflows
di: Chen, June, et al.
Pubblicazione: (2026)
di: Chen, June, et al.
Pubblicazione: (2026)
Malleus: Straggler-Resilient Hybrid Parallel Training of Large-scale Models via Malleable Data and Model Parallelization
di: Li, Haoyang, et al.
Pubblicazione: (2024)
di: Li, Haoyang, et al.
Pubblicazione: (2024)
AntDT: A Self-Adaptive Distributed Training Framework for Leader and Straggler Nodes
di: Xiao, Youshao, et al.
Pubblicazione: (2024)
di: Xiao, Youshao, et al.
Pubblicazione: (2024)
Cooperative Inference with Interleaved Operator Partitioning for CNNs
di: Liu, Zhibang, et al.
Pubblicazione: (2024)
di: Liu, Zhibang, et al.
Pubblicazione: (2024)
Efficient CPU-GPU Collaborative Inference for MoE-based LLMs on Memory-Limited Systems
di: Huang, En-Ming, et al.
Pubblicazione: (2025)
di: Huang, En-Ming, et al.
Pubblicazione: (2025)
Uncoded Storage Coded Transmission Elastic Computing with Straggler Tolerance in Heterogeneous Systems
di: Zhong, Xi, et al.
Pubblicazione: (2024)
di: Zhong, Xi, et al.
Pubblicazione: (2024)
Distributed Inference Performance Optimization for LLMs on CPUs
di: He, Pujiang, et al.
Pubblicazione: (2024)
di: He, Pujiang, et al.
Pubblicazione: (2024)
OOCO: Latency-disaggregated Architecture for Online-Offline Co-locate LLM Serving
di: Wu, Siyu, et al.
Pubblicazione: (2025)
di: Wu, Siyu, et al.
Pubblicazione: (2025)
Understanding Stragglers in Large Model Training Using What-if Analysis
di: Lin, Jinkun, et al.
Pubblicazione: (2025)
di: Lin, Jinkun, et al.
Pubblicazione: (2025)
COPUS: Co-adaptive Parallelism and Batch Size Selection in Large Language Model Training
di: Sakip, Akhmed, et al.
Pubblicazione: (2026)
di: Sakip, Akhmed, et al.
Pubblicazione: (2026)
TSUE: A Two-Stage Data Update Method for an Erasure Coded Cluster File System
di: Wei, Zheng, et al.
Pubblicazione: (2025)
di: Wei, Zheng, et al.
Pubblicazione: (2025)
Ensuring Data Privacy in AC Optimal Power Flow with a Distributed Co-Simulation Framework
di: Dai, Xinliang, et al.
Pubblicazione: (2024)
di: Dai, Xinliang, et al.
Pubblicazione: (2024)
Lightweight Federated Learning with Differential Privacy and Straggler Resilience
di: Hong, Shu, et al.
Pubblicazione: (2024)
di: Hong, Shu, et al.
Pubblicazione: (2024)
Orchestrated Co-scheduling, Resource Partitioning, and Power Capping on CPU-GPU Heterogeneous Systems via Machine Learning
di: Saba, Issa, et al.
Pubblicazione: (2024)
di: Saba, Issa, et al.
Pubblicazione: (2024)
Designing Co-operation in Systems of Hierarchical, Multi-objective Schedulers for Stream Processing
di: Dangwal, Animesh, et al.
Pubblicazione: (2025)
di: Dangwal, Animesh, et al.
Pubblicazione: (2025)
AdaBridge: Dynamic Data and Computation Reuse for Efficient Multi-task DNN Co-evolution in Edge Systems
di: Wang, Lehao, et al.
Pubblicazione: (2024)
di: Wang, Lehao, et al.
Pubblicazione: (2024)
Straggler-Resilient Decentralized Learning via Adaptive Asynchronous Updates
di: Xiong, Guojun, et al.
Pubblicazione: (2023)
di: Xiong, Guojun, et al.
Pubblicazione: (2023)
Documenti analoghi
-
Exploiting Stragglers in Distributed Computing Systems with Task Grouping
di: Adikari, Tharindu, et al.
Pubblicazione: (2024) -
FlowSpec: Continuous Pipelined Speculative Decoding for Efficient Distributed LLM Inference
di: Liu, Xing, et al.
Pubblicazione: (2025) -
dSTAR: Straggler Tolerant and Byzantine Resilient Distributed SGD
di: Yan, Jiahe, et al.
Pubblicazione: (2024) -
Sparsity-Preserving Encodings for Straggler-Optimal Distributed Matrix Computations at the Edge
di: Das, Anindya Bijoy, et al.
Pubblicazione: (2024) -
Distributed Learning based on 1-Bit Gradient Coding in the Presence of Stragglers
di: Li, Chengxi, et al.
Pubblicazione: (2024)