HyperParallel: A Supernode-Affinity AI Framework
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Xin, Sun, Beilei, Su, Teng, Zhang, Qinghua, Bao, Chong, Chen, Lei, Jin, Xuefeng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
HyperOffload: Graph-Driven Hierarchical Memory Management for Large Language Models on SuperNode Architectures
von: Liu, Fangxin, et al.
Veröffentlicht: (2026)
von: Liu, Fangxin, et al.
Veröffentlicht: (2026)
ICPS: Real-Time Resource Configuration for Cloud Serverless Functions Considering Affinity
von: Chen, Long, et al.
Veröffentlicht: (2025)
von: Chen, Long, et al.
Veröffentlicht: (2025)
InternEvo: Efficient Long-sequence Large Language Model Training via Hybrid Parallelism and Redundant Sharding
von: Chen, Qiaoling, et al.
Veröffentlicht: (2024)
von: Chen, Qiaoling, et al.
Veröffentlicht: (2024)
Amoeba: Runtime Tensor Parallel Transformation for LLM Inference Services
von: Chen, Haoyu, et al.
Veröffentlicht: (2025)
von: Chen, Haoyu, et al.
Veröffentlicht: (2025)
LoongTrain: Efficient Training of Long-Sequence LLMs with Head-Context Parallelism
von: Gu, Diandian, et al.
Veröffentlicht: (2024)
von: Gu, Diandian, et al.
Veröffentlicht: (2024)
DawnPiper: A Memory-scablable Pipeline Parallel Training Framework
von: Peng, Xuan, et al.
Veröffentlicht: (2025)
von: Peng, Xuan, et al.
Veröffentlicht: (2025)
ZeroPP: Unleashing Exceptional Parallelism Efficiency through Tensor-Parallelism-Free Methodology
von: Tang, Ding, et al.
Veröffentlicht: (2024)
von: Tang, Ding, et al.
Veröffentlicht: (2024)
Affinity-aware Serverless Function Scheduling
von: De Palma, Giuseppe, et al.
Veröffentlicht: (2024)
von: De Palma, Giuseppe, et al.
Veröffentlicht: (2024)
KCES: A Workflow Containerization Scheduling Scheme Under Cloud-Edge Collaboration Framework
von: Shan, Chenggang, et al.
Veröffentlicht: (2024)
von: Shan, Chenggang, et al.
Veröffentlicht: (2024)
MegaScale-Omni: A Hyper-Scale, Workload-Resilient System for MultiModal LLM Training in Production
von: Xue, Chunyu, et al.
Veröffentlicht: (2026)
von: Xue, Chunyu, et al.
Veröffentlicht: (2026)
SPPO:Efficient Long-sequence LLM Training via Adaptive Sequence Pipeline Parallel Offloading
von: Chen, Qiaoling, et al.
Veröffentlicht: (2025)
von: Chen, Qiaoling, et al.
Veröffentlicht: (2025)
HAP: Hybrid Adaptive Parallelism for Efficient Mixture-of-Experts Inference
von: Lin, Haoran, et al.
Veröffentlicht: (2025)
von: Lin, Haoran, et al.
Veröffentlicht: (2025)
Surviving Partial Rank Failures in Wide Expert-Parallel MoE Inference
von: Sun, Xun, et al.
Veröffentlicht: (2026)
von: Sun, Xun, et al.
Veröffentlicht: (2026)
SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference
von: Zhao, Yihao, et al.
Veröffentlicht: (2025)
von: Zhao, Yihao, et al.
Veröffentlicht: (2025)
NeutronTP: Load-Balanced Distributed Full-Graph GNN Training with Tensor Parallelism
von: Ai, Xin, et al.
Veröffentlicht: (2024)
von: Ai, Xin, et al.
Veröffentlicht: (2024)
Resource-efficient Parallel Split Learning in Heterogeneous Edge Computing
von: Zhang, Mingjin, et al.
Veröffentlicht: (2024)
von: Zhang, Mingjin, et al.
Veröffentlicht: (2024)
DualMap: Enabling Both Cache Affinity and Load Balancing for Distributed LLM Serving
von: Yuan, Ying, et al.
Veröffentlicht: (2026)
von: Yuan, Ying, et al.
Veröffentlicht: (2026)
NanoCP: Request-Level Dynamic Context Parallelism for Data-Expert Parallel Decoding
von: Chen, Jiefei, et al.
Veröffentlicht: (2026)
von: Chen, Jiefei, et al.
Veröffentlicht: (2026)
Communication-Computation Pipeline Parallel Split Learning over Wireless Edge Networks
von: Liu, Chenyu, et al.
Veröffentlicht: (2025)
von: Liu, Chenyu, et al.
Veröffentlicht: (2025)
Resource Allocation in HyperX Networks
von: Cano, Alejandro, et al.
Veröffentlicht: (2026)
von: Cano, Alejandro, et al.
Veröffentlicht: (2026)
Synergistic Tensor and Pipeline Parallelism
von: Qi, Mengshi, et al.
Veröffentlicht: (2025)
von: Qi, Mengshi, et al.
Veröffentlicht: (2025)
Seq1F1B: Efficient Sequence-Level Pipeline Parallelism for Large Language Model Training
von: Sun, Ao, et al.
Veröffentlicht: (2024)
von: Sun, Ao, et al.
Veröffentlicht: (2024)
MoFa: A Unified Performance Modeling Framework for LLM Pretraining
von: Zhao, Lu, et al.
Veröffentlicht: (2025)
von: Zhao, Lu, et al.
Veröffentlicht: (2025)
Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism
von: Qing, Yuhao, et al.
Veröffentlicht: (2025)
von: Qing, Yuhao, et al.
Veröffentlicht: (2025)
UniPar: A Unified LLM-Based Framework for Parallel and Accelerated Code Translation in HPC
von: Bitan, Tomer, et al.
Veröffentlicht: (2025)
von: Bitan, Tomer, et al.
Veröffentlicht: (2025)
Heterogeneous Federated Fine-Tuning with Parallel One-Rank Adaptation
von: Zhang, Zikai, et al.
Veröffentlicht: (2026)
von: Zhang, Zikai, et al.
Veröffentlicht: (2026)
A Flexible Programmable Pipeline Parallelism Framework for Efficient DNN Training
von: Jiang, Lijuan, et al.
Veröffentlicht: (2025)
von: Jiang, Lijuan, et al.
Veröffentlicht: (2025)
Communication-Efficient Serving for Video Diffusion Models with Latent Parallelism
von: Wu, Zhiyuan, et al.
Veröffentlicht: (2025)
von: Wu, Zhiyuan, et al.
Veröffentlicht: (2025)
CFP: Efficient Optimization of Intra-Operator Parallelism Plans for Large Model Training
von: Hu, Weifang, et al.
Veröffentlicht: (2025)
von: Hu, Weifang, et al.
Veröffentlicht: (2025)
ElasWave: An Elastic-Native System for Scalable Hybrid-Parallel Training
von: Kang, Xueze, et al.
Veröffentlicht: (2025)
von: Kang, Xueze, et al.
Veröffentlicht: (2025)
Power Aware Container Placement in Cloud Computing with Affinity and Cubic Power Model
von: Sarkar, Suvarthi, et al.
Veröffentlicht: (2024)
von: Sarkar, Suvarthi, et al.
Veröffentlicht: (2024)
FastSet: Parallel Claim Settlement
von: Chen, Xiaohong, et al.
Veröffentlicht: (2025)
von: Chen, Xiaohong, et al.
Veröffentlicht: (2025)
Balancing Pipeline Parallelism with Vocabulary Parallelism
von: Yeung, Man Tsung, et al.
Veröffentlicht: (2024)
von: Yeung, Man Tsung, et al.
Veröffentlicht: (2024)
Joint Temporal-Structural Representation Learning for Distributed Fault Discrimination in Microservice Architectures
von: Xue, Yihan, et al.
Veröffentlicht: (2026)
von: Xue, Yihan, et al.
Veröffentlicht: (2026)
Mapping Parallel Matrix Multiplication in GotoBLAS2 to the AMD Versal ACAP for Deep Learning
von: Lei, Jie, et al.
Veröffentlicht: (2024)
von: Lei, Jie, et al.
Veröffentlicht: (2024)
TD-Pipe: Temporally-Disaggregated Pipeline Parallelism Architecture for High-Throughput LLM Inference
von: Zhang, Hongbin, et al.
Veröffentlicht: (2025)
von: Zhang, Hongbin, et al.
Veröffentlicht: (2025)
Optimizing Long-context LLM Serving via Fine-grained Sequence Parallelism
von: Li, Cong, et al.
Veröffentlicht: (2025)
von: Li, Cong, et al.
Veröffentlicht: (2025)
Accelerating Microswimmer Simulations via a Heterogeneous Pipelined Parallel-in-Time Framework
von: Huang, Ruixiang, et al.
Veröffentlicht: (2026)
von: Huang, Ruixiang, et al.
Veröffentlicht: (2026)
Speeding up Local Optimization in Vehicle Routing with Tensor-based GPU Acceleration
von: Lei, Zhenyu, et al.
Veröffentlicht: (2025)
von: Lei, Zhenyu, et al.
Veröffentlicht: (2025)
gLLM: Global Balanced Pipeline Parallelism System for Distributed LLM Serving with Token Throttling
von: Guo, Tianyu, et al.
Veröffentlicht: (2025)
von: Guo, Tianyu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
HyperOffload: Graph-Driven Hierarchical Memory Management for Large Language Models on SuperNode Architectures
von: Liu, Fangxin, et al.
Veröffentlicht: (2026) -
ICPS: Real-Time Resource Configuration for Cloud Serverless Functions Considering Affinity
von: Chen, Long, et al.
Veröffentlicht: (2025) -
InternEvo: Efficient Long-sequence Large Language Model Training via Hybrid Parallelism and Redundant Sharding
von: Chen, Qiaoling, et al.
Veröffentlicht: (2024) -
Amoeba: Runtime Tensor Parallel Transformation for LLM Inference Services
von: Chen, Haoyu, et al.
Veröffentlicht: (2025) -
LoongTrain: Efficient Training of Long-Sequence LLMs with Head-Context Parallelism
von: Gu, Diandian, et al.
Veröffentlicht: (2024)