HetCCL: Accelerating LLM Training with Heterogeneous GPUs
Fuente:
arXiv
Salvato in:
| Autori principali: | Kim, Heehoon, Lee, Jaehwan, Kim, Taejeoung, Park, Jongwon, Kim, Jinpyo, Suh, Pyongwon, Choi, Ryan H., Lee, Sangwoo, Lee, Jaejin |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Accelerating Sparse Matrix-Matrix Multiplication on GPUs with Processing Near HBMs
di: Li, Shiju, et al.
Pubblicazione: (2025)
di: Li, Shiju, et al.
Pubblicazione: (2025)
DFLOP: A Data-driven Framework for Multimodal LLM Training Pipeline Optimization
di: An, Hyeonjun, et al.
Pubblicazione: (2026)
di: An, Hyeonjun, et al.
Pubblicazione: (2026)
AGAThA: Fast and Efficient GPU Acceleration of Guided Sequence Alignment for Long Read Mapping
di: Park, Seongyeon, et al.
Pubblicazione: (2024)
di: Park, Seongyeon, et al.
Pubblicazione: (2024)
PIM-SHERPA: Software Method for On-device LLM Inference by Resolving PIM Memory Attribute and Layout Inconsistencies
di: Lee, Sunjung, et al.
Pubblicazione: (2026)
di: Lee, Sunjung, et al.
Pubblicazione: (2026)
Characterizing Compute-Communication Overlap in GPU-Accelerated Distributed Deep Learning: Performance and Power Implications
di: Lee, Seonho, et al.
Pubblicazione: (2025)
di: Lee, Seonho, et al.
Pubblicazione: (2025)
Demystifying Cost-Efficiency in LLM Serving over Heterogeneous GPUs
di: Jiang, Youhe, et al.
Pubblicazione: (2025)
di: Jiang, Youhe, et al.
Pubblicazione: (2025)
HetGPU: The pursuit of making binary compatibility towards GPUs
di: Yang, Yiwei, et al.
Pubblicazione: (2025)
di: Yang, Yiwei, et al.
Pubblicazione: (2025)
HetRL: Efficient Reinforcement Learning for LLMs in Heterogeneous Environments
di: He, Yongjun, et al.
Pubblicazione: (2025)
di: He, Yongjun, et al.
Pubblicazione: (2025)
ZipCCL: Efficient Lossless Data Compression of Communication Collectives for Accelerating LLM Training
di: Lin, Wenxiang, et al.
Pubblicazione: (2026)
di: Lin, Wenxiang, et al.
Pubblicazione: (2026)
Ding-Dong Ditch: Peeking Into Spot Instance Availability
di: Kim, Kyumin, et al.
Pubblicazione: (2026)
di: Kim, Kyumin, et al.
Pubblicazione: (2026)
FREESH: Fair, Resource- and Energy-Efficient Scheduling for LLM Serving on Heterogeneous GPUs
di: He, Xuan, et al.
Pubblicazione: (2025)
di: He, Xuan, et al.
Pubblicazione: (2025)
PointSplit: Towards On-device 3D Object Detection with Heterogeneous Low-power Accelerators
di: Park, Keondo, et al.
Pubblicazione: (2025)
di: Park, Keondo, et al.
Pubblicazione: (2025)
FlashSketch: Sketch-Kernel Co-Design for Fast Sparse Sketching on GPUs
di: Dwaraknath, Rajat Vadiraj, et al.
Pubblicazione: (2026)
di: Dwaraknath, Rajat Vadiraj, et al.
Pubblicazione: (2026)
SpotVista: Availability-Aware Recommendation System for Reliable and Cost-Efficient Multi-Node Spot Instances
di: Kim, Taeyoon, et al.
Pubblicazione: (2026)
di: Kim, Taeyoon, et al.
Pubblicazione: (2026)
Accelerating Maximal Biclique Enumeration on GPUs
di: Hsieh, Chou-Ying, et al.
Pubblicazione: (2024)
di: Hsieh, Chou-Ying, et al.
Pubblicazione: (2024)
LUT-GEMM: Quantized Matrix Multiplication based on LUTs for Efficient Inference in Large-Scale Generative Language Models
di: Park, Gunho, et al.
Pubblicazione: (2022)
di: Park, Gunho, et al.
Pubblicazione: (2022)
Efficient LLM Inference with Activation Checkpointing and Hybrid Caching
di: Lee, Sanghyeon, et al.
Pubblicazione: (2025)
di: Lee, Sanghyeon, et al.
Pubblicazione: (2025)
PrefillShare: A Shared Prefill Module for KV Reuse in Multi-LLM Disaggregated Serving
di: Woo, Sunghyeon, et al.
Pubblicazione: (2026)
di: Woo, Sunghyeon, et al.
Pubblicazione: (2026)
BOute: Cost-Efficient LLM Serving with Heterogeneous LLMs and GPUs via Multi-Objective Bayesian Optimization
di: Jiang, Youhe, et al.
Pubblicazione: (2026)
di: Jiang, Youhe, et al.
Pubblicazione: (2026)
A Case Study of API Design for Interoperability and Security of the Internet of Things
di: Kim, Dongha, et al.
Pubblicazione: (2024)
di: Kim, Dongha, et al.
Pubblicazione: (2024)
KubePACS: Kubernetes Cluster Using Performant, Highly Available, and Cost Efficient Spot Instances
di: Kim, Taeyoon, et al.
Pubblicazione: (2026)
di: Kim, Taeyoon, et al.
Pubblicazione: (2026)
Breaking the Capacity Bottleneck in Model-Heterogeneous Federated Learning via Gradual Model Restoration
di: Ma, Chengjie, et al.
Pubblicazione: (2025)
di: Ma, Chengjie, et al.
Pubblicazione: (2025)
FlexiWalker: Extensible GPU Framework for Efficient Dynamic Random Walks with Runtime Adaptation
di: Park, Seongyeon, et al.
Pubblicazione: (2025)
di: Park, Seongyeon, et al.
Pubblicazione: (2025)
Multi-Resolution Model Fusion for Accelerating the Convolutional Neural Network Training
di: Wang, Kewei, et al.
Pubblicazione: (2025)
di: Wang, Kewei, et al.
Pubblicazione: (2025)
ExeGPT: Constraint-Aware Resource Scheduling for LLM Inference
di: Oh, Hyungjun, et al.
Pubblicazione: (2024)
di: Oh, Hyungjun, et al.
Pubblicazione: (2024)
Accelerating Particle-Mesh Algorithms with FPGAs and OmpSs@OpenCL
di: Guidotti, Nicolas Lee
Pubblicazione: (2025)
di: Guidotti, Nicolas Lee
Pubblicazione: (2025)
GPU-Accelerated Distributed QAOA on Large-scale HPC Ecosystems
di: Xu, Zhihao, et al.
Pubblicazione: (2025)
di: Xu, Zhihao, et al.
Pubblicazione: (2025)
Straggler Tolerant and Resilient DL Training on Homogeneous GPUs
di: Zhang, Zeyu, et al.
Pubblicazione: (2025)
di: Zhang, Zeyu, et al.
Pubblicazione: (2025)
Accelerating LLM Inference with Precomputed Query Storage
di: Park, Jay H., et al.
Pubblicazione: (2025)
di: Park, Jay H., et al.
Pubblicazione: (2025)
From Tokens to Layers: Redefining Stall-Free Scheduling for MoE Serving with Layered Prefill
di: Lee, Gunjun, et al.
Pubblicazione: (2025)
di: Lee, Gunjun, et al.
Pubblicazione: (2025)
Why Do AI Agents Systematically Fail at Cloud Root Cause Analysis?
di: Kim, Taeyoon, et al.
Pubblicazione: (2026)
di: Kim, Taeyoon, et al.
Pubblicazione: (2026)
Heterogeneous Federated Learning with Prototype Alignment and Upscaling
di: Lee, Gyuejeong, et al.
Pubblicazione: (2025)
di: Lee, Gyuejeong, et al.
Pubblicazione: (2025)
Accelerating Compound LLM Training Workloads with Maestro
di: Yuan, Xiulong, et al.
Pubblicazione: (2026)
di: Yuan, Xiulong, et al.
Pubblicazione: (2026)
Co-LoRA: Collaborative Model Personalization on Heterogeneous Multi-Modal Clients
di: Seo, Minhyuk, et al.
Pubblicazione: (2025)
di: Seo, Minhyuk, et al.
Pubblicazione: (2025)
AMPED: Accelerating MTTKRP for Billion-Scale Sparse Tensor Decomposition on Multiple GPUs
di: Wijeratne, Sasindu, et al.
Pubblicazione: (2025)
di: Wijeratne, Sasindu, et al.
Pubblicazione: (2025)
Popcorn: Accelerating Kernel K-means on GPUs through Sparse Linear Algebra
di: Bellavita, Julian, et al.
Pubblicazione: (2025)
di: Bellavita, Julian, et al.
Pubblicazione: (2025)
Accelerating high-order continuum kinetic plasma simulations using multiple GPUs
di: Ho, Andrew, et al.
Pubblicazione: (2024)
di: Ho, Andrew, et al.
Pubblicazione: (2024)
TrioSeq: A Novel Approach to Accelerate Triplet Sequence Alignment on GPUs
di: Graça, Miguel, et al.
Pubblicazione: (2026)
di: Graça, Miguel, et al.
Pubblicazione: (2026)
BlockLLM: Multi-tenant Finer-grained Serving for Large Language Models
di: Hu, Bodun, et al.
Pubblicazione: (2024)
di: Hu, Bodun, et al.
Pubblicazione: (2024)
HeterMoE: Efficient Training of Mixture-of-Experts Models on Heterogeneous GPUs
di: Wu, Yongji, et al.
Pubblicazione: (2025)
di: Wu, Yongji, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Accelerating Sparse Matrix-Matrix Multiplication on GPUs with Processing Near HBMs
di: Li, Shiju, et al.
Pubblicazione: (2025) -
DFLOP: A Data-driven Framework for Multimodal LLM Training Pipeline Optimization
di: An, Hyeonjun, et al.
Pubblicazione: (2026) -
AGAThA: Fast and Efficient GPU Acceleration of Guided Sequence Alignment for Long Read Mapping
di: Park, Seongyeon, et al.
Pubblicazione: (2024) -
PIM-SHERPA: Software Method for On-device LLM Inference by Resolving PIM Memory Attribute and Layout Inconsistencies
di: Lee, Sunjung, et al.
Pubblicazione: (2026) -
Characterizing Compute-Communication Overlap in GPU-Accelerated Distributed Deep Learning: Performance and Power Implications
di: Lee, Seonho, et al.
Pubblicazione: (2025)