NanoCP: Request-Level Dynamic Context Parallelism for Data-Expert Parallel Decoding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chen, Jiefei, Lin, Binbin, Ma, Jinming, Duan, Jiangfei, Duanmu, Haojie, Liu, Hao, Cheng, Qinxiu, Li, Xiuhong, Pei, Zhilin, Wang, Hui, Zhang, Xingcheng, Lin, Dahua |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MuxServe: Flexible Spatial-Temporal Multiplexing for Multiple LLM Serving
von: Duan, Jiangfei, et al.
Veröffentlicht: (2024)
von: Duan, Jiangfei, et al.
Veröffentlicht: (2024)
ZeroPP: Unleashing Exceptional Parallelism Efficiency through Tensor-Parallelism-Free Methodology
von: Tang, Ding, et al.
Veröffentlicht: (2024)
von: Tang, Ding, et al.
Veröffentlicht: (2024)
Zeppelin: Balancing Variable-length Workloads in Data Parallel Large Model Training
von: Chen, Chang, et al.
Veröffentlicht: (2025)
von: Chen, Chang, et al.
Veröffentlicht: (2025)
Parcae: Proactive, Liveput-Optimized DNN Training on Preemptible Instances
von: Duan, Jiangfei, et al.
Veröffentlicht: (2024)
von: Duan, Jiangfei, et al.
Veröffentlicht: (2024)
Nitsum: Serving Tiered LLM Requests with Adaptive Tensor Parallelism
von: Srivatsa, Vikranth, et al.
Veröffentlicht: (2026)
von: Srivatsa, Vikranth, et al.
Veröffentlicht: (2026)
Balancing Pipeline Parallelism with Vocabulary Parallelism
von: Yeung, Man Tsung, et al.
Veröffentlicht: (2024)
von: Yeung, Man Tsung, et al.
Veröffentlicht: (2024)
ResiHP: Taming LLM Training Failures with Dynamic Hybrid Parallelism
von: Ma, Tenghui, et al.
Veröffentlicht: (2026)
von: Ma, Tenghui, et al.
Veröffentlicht: (2026)
HAP: Hybrid Adaptive Parallelism for Efficient Mixture-of-Experts Inference
von: Lin, Haoran, et al.
Veröffentlicht: (2025)
von: Lin, Haoran, et al.
Veröffentlicht: (2025)
TASP: Topology-aware Sequence Parallelism
von: Wang, Yida, et al.
Veröffentlicht: (2025)
von: Wang, Yida, et al.
Veröffentlicht: (2025)
Fold-CP: A Context Parallelism Framework for Biomolecular Modeling
von: Lin, Dejun, et al.
Veröffentlicht: (2026)
von: Lin, Dejun, et al.
Veröffentlicht: (2026)
Surviving Partial Rank Failures in Wide Expert-Parallel MoE Inference
von: Sun, Xun, et al.
Veröffentlicht: (2026)
von: Sun, Xun, et al.
Veröffentlicht: (2026)
Efficient Training of Large Language Models on Distributed Infrastructures: A Survey
von: Duan, Jiangfei, et al.
Veröffentlicht: (2024)
von: Duan, Jiangfei, et al.
Veröffentlicht: (2024)
SKVQ: Sliding-window Key and Value Cache Quantization for Large Language Models
von: Duanmu, Haojie, et al.
Veröffentlicht: (2024)
von: Duanmu, Haojie, et al.
Veröffentlicht: (2024)
AdaPtis: Reducing Pipeline Bubbles with Adaptive Pipeline Parallelism on Heterogeneous Models
von: Guo, Jihu, et al.
Veröffentlicht: (2025)
von: Guo, Jihu, et al.
Veröffentlicht: (2025)
MegaScale-Infer: Serving Mixture-of-Experts at Scale with Disaggregated Expert Parallelism
von: Zhu, Ruidong, et al.
Veröffentlicht: (2025)
von: Zhu, Ruidong, et al.
Veröffentlicht: (2025)
Efficient MoE Inference with Fine-Grained Scheduling of Disaggregated Expert Parallelism
von: Pan, Xinglin, et al.
Veröffentlicht: (2025)
von: Pan, Xinglin, et al.
Veröffentlicht: (2025)
Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism
von: Wei, Jinhui, et al.
Veröffentlicht: (2025)
von: Wei, Jinhui, et al.
Veröffentlicht: (2025)
Malleus: Straggler-Resilient Hybrid Parallel Training of Large-scale Models via Malleable Data and Model Parallelization
von: Li, Haoyang, et al.
Veröffentlicht: (2024)
von: Li, Haoyang, et al.
Veröffentlicht: (2024)
Enhancing Memory Efficiency in Large Language Model Training Through Chronos-aware Pipeline Parallelism
von: Lin, Xinyuan, et al.
Veröffentlicht: (2025)
von: Lin, Xinyuan, et al.
Veröffentlicht: (2025)
UniEP: Unified Expert-Parallel MoE MegaKernel for LLM Training
von: Zheng, Size, et al.
Veröffentlicht: (2026)
von: Zheng, Size, et al.
Veröffentlicht: (2026)
Concurrent Scheduling of High-Level Parallel Programs on Multi-GPU Systems
von: Knorr, Fabian, et al.
Veröffentlicht: (2025)
von: Knorr, Fabian, et al.
Veröffentlicht: (2025)
H2:Towards Efficient Large-Scale LLM Training on Hyper-Heterogeneous Cluster over 1,000 Chips
von: Tang, Ding, et al.
Veröffentlicht: (2025)
von: Tang, Ding, et al.
Veröffentlicht: (2025)
Unleashing Scalable Context Parallelism for Foundation Models Pre-Training via FCP
von: Zhao, Yilong, et al.
Veröffentlicht: (2026)
von: Zhao, Yilong, et al.
Veröffentlicht: (2026)
LoongTrain: Efficient Training of Long-Sequence LLMs with Head-Context Parallelism
von: Gu, Diandian, et al.
Veröffentlicht: (2024)
von: Gu, Diandian, et al.
Veröffentlicht: (2024)
MoEntwine: Unleashing the Potential of Wafer-scale Chips for Large-scale Expert Parallel Inference
von: Tang, Xinru, et al.
Veröffentlicht: (2025)
von: Tang, Xinru, et al.
Veröffentlicht: (2025)
The Entropy of Parallel Systems
von: Adefemi, Temitayo
Veröffentlicht: (2025)
von: Adefemi, Temitayo
Veröffentlicht: (2025)
Lectures on Parallel Computing
von: Träff, Jesper Larsson
Veröffentlicht: (2024)
von: Träff, Jesper Larsson
Veröffentlicht: (2024)
APEX: An Extensible and Dynamism-Aware Simulator for Automated Parallel Execution in LLM Serving
von: Lin, Yi-Chien, et al.
Veröffentlicht: (2024)
von: Lin, Yi-Chien, et al.
Veröffentlicht: (2024)
Shortcut-connected Expert Parallelism for Accelerating Mixture-of-Experts
von: Cai, Weilin, et al.
Veröffentlicht: (2024)
von: Cai, Weilin, et al.
Veröffentlicht: (2024)
SparseServe: Unlocking Parallelism for Dynamic Sparse Attention in Long-Context LLM Serving
von: Zhou, Qihui, et al.
Veröffentlicht: (2025)
von: Zhou, Qihui, et al.
Veröffentlicht: (2025)
Adaptra: Straggler-Resilient Hybrid-Parallel Training with Pipeline Adaptation
von: Wu, Tianyuan, et al.
Veröffentlicht: (2025)
von: Wu, Tianyuan, et al.
Veröffentlicht: (2025)
Synergistic Tensor and Pipeline Parallelism
von: Qi, Mengshi, et al.
Veröffentlicht: (2025)
von: Qi, Mengshi, et al.
Veröffentlicht: (2025)
OMP-Engineer: Bridging Syntax Analysis and In-Context Learning for Efficient Automated OpenMP Parallelization
von: Wang, Weidong, et al.
Veröffentlicht: (2024)
von: Wang, Weidong, et al.
Veröffentlicht: (2024)
A Survey on Parallel Text Generation: From Parallel Decoding to Diffusion Language Models
von: Zhang, Lingzhe, et al.
Veröffentlicht: (2025)
von: Zhang, Lingzhe, et al.
Veröffentlicht: (2025)
FALCON: Pinpointing and Mitigating Stragglers for Large-Scale Hybrid-Parallel Training
von: Wu, Tianyuan, et al.
Veröffentlicht: (2024)
von: Wu, Tianyuan, et al.
Veröffentlicht: (2024)
StarTrail: Concentric Ring Sequence Parallelism for Efficient Near-Infinite-Context Transformer Model Training
von: Liu, Ziming, et al.
Veröffentlicht: (2024)
von: Liu, Ziming, et al.
Veröffentlicht: (2024)
FastSet: Parallel Claim Settlement
von: Chen, Xiaohong, et al.
Veröffentlicht: (2025)
von: Chen, Xiaohong, et al.
Veröffentlicht: (2025)
Parallelizing Maximal Clique Enumeration on GPUs
von: Almasri, Mohammad, et al.
Veröffentlicht: (2022)
von: Almasri, Mohammad, et al.
Veröffentlicht: (2022)
Zero Bubble Pipeline Parallelism
von: Qi, Penghui, et al.
Veröffentlicht: (2023)
von: Qi, Penghui, et al.
Veröffentlicht: (2023)
Pipeline Parallelism with Controllable Memory
von: Qi, Penghui, et al.
Veröffentlicht: (2024)
von: Qi, Penghui, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
MuxServe: Flexible Spatial-Temporal Multiplexing for Multiple LLM Serving
von: Duan, Jiangfei, et al.
Veröffentlicht: (2024) -
ZeroPP: Unleashing Exceptional Parallelism Efficiency through Tensor-Parallelism-Free Methodology
von: Tang, Ding, et al.
Veröffentlicht: (2024) -
Zeppelin: Balancing Variable-length Workloads in Data Parallel Large Model Training
von: Chen, Chang, et al.
Veröffentlicht: (2025) -
Parcae: Proactive, Liveput-Optimized DNN Training on Preemptible Instances
von: Duan, Jiangfei, et al.
Veröffentlicht: (2024) -
Nitsum: Serving Tiered LLM Requests with Adaptive Tensor Parallelism
von: Srivatsa, Vikranth, et al.
Veröffentlicht: (2026)