GOGH: Correlation-Guided Orchestration of GPUs in Heterogeneous Clusters
Fuente:
arXiv
Salvato in:
| Autori principali: | Raeisi, Ahmad, Dolati, Mahdi, Darabi, Sina, Talebi, Sadegh, Eugster, Patrick, Khonsari, Ahmad |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
HetCCL: Accelerating LLM Training with Heterogeneous GPUs
di: Kim, Heehoon, et al.
Pubblicazione: (2026)
di: Kim, Heehoon, et al.
Pubblicazione: (2026)
Tessera: Unlocking Heterogeneous GPUs through Kernel-Granularity Disaggregation
di: Hu, Tiancheng, et al.
Pubblicazione: (2026)
di: Hu, Tiancheng, et al.
Pubblicazione: (2026)
HeterMoE: Efficient Training of Mixture-of-Experts Models on Heterogeneous GPUs
di: Wu, Yongji, et al.
Pubblicazione: (2025)
di: Wu, Yongji, et al.
Pubblicazione: (2025)
AReaL-Hex: Accommodating Asynchronous RL Training over Heterogeneous GPUs
di: Yan, Ran, et al.
Pubblicazione: (2025)
di: Yan, Ran, et al.
Pubblicazione: (2025)
PlexRL: Cluster-Level Orchestration of Serviceized LLM Execution for RLVR
di: Zhang, Yiqi, et al.
Pubblicazione: (2026)
di: Zhang, Yiqi, et al.
Pubblicazione: (2026)
LLMQ: Efficient Lower-Precision Pretraining for Consumer GPUs
di: Schultheis, Erik, et al.
Pubblicazione: (2025)
di: Schultheis, Erik, et al.
Pubblicazione: (2025)
Communication-Avoiding Linear Algebraic Kernel K-Means on GPUs
di: Bellavita, Julian, et al.
Pubblicazione: (2026)
di: Bellavita, Julian, et al.
Pubblicazione: (2026)
Scaling State-Space Models on Multiple GPUs with Tensor Parallelism
di: Dutt, Anurag, et al.
Pubblicazione: (2026)
di: Dutt, Anurag, et al.
Pubblicazione: (2026)
AcceleratedLiNGAM: Learning Causal DAGs at the speed of GPUs
di: Akinwande, Victor, et al.
Pubblicazione: (2024)
di: Akinwande, Victor, et al.
Pubblicazione: (2024)
Federated Incomplete Multi-View Clustering with Heterogeneous Graph Neural Networks
di: Yan, Xueming, et al.
Pubblicazione: (2024)
di: Yan, Xueming, et al.
Pubblicazione: (2024)
Helix: Serving Large Language Models over Heterogeneous GPUs and Network via Max-Flow
di: Mei, Yixuan, et al.
Pubblicazione: (2024)
di: Mei, Yixuan, et al.
Pubblicazione: (2024)
FLUX: Fast Software-based Communication Overlap On GPUs Through Kernel Fusion
di: Chang, Li-Wen, et al.
Pubblicazione: (2024)
di: Chang, Li-Wen, et al.
Pubblicazione: (2024)
PipeFill: Using GPUs During Bubbles in Pipeline-parallel LLM Training
di: Arfeen, Daiyaan, et al.
Pubblicazione: (2024)
di: Arfeen, Daiyaan, et al.
Pubblicazione: (2024)
Parallel Split Learning with Global Sampling
di: Kohankhaki, Mohammad, et al.
Pubblicazione: (2024)
di: Kohankhaki, Mohammad, et al.
Pubblicazione: (2024)
Metadata-Guided Adaptable Frequency Scaling across Heterogeneous Applications and Devices
di: Yan, Jinqi, et al.
Pubblicazione: (2025)
di: Yan, Jinqi, et al.
Pubblicazione: (2025)
HARP: Orchestrating Automated Parallel Training on Heterogeneous GPU Clusters
di: Liang, Antian, et al.
Pubblicazione: (2025)
di: Liang, Antian, et al.
Pubblicazione: (2025)
When GPUs Fail Quietly: Observability-Aware Early Warning Beyond Numeric Telemetry
di: Bidollahkhani, Michael, et al.
Pubblicazione: (2026)
di: Bidollahkhani, Michael, et al.
Pubblicazione: (2026)
FedORGP: Guiding Heterogeneous Federated Learning with Orthogonality Regularization on Global Prototypes
di: Guo, Fucheng, et al.
Pubblicazione: (2025)
di: Guo, Fucheng, et al.
Pubblicazione: (2025)
Catastrophic Forgetting Resilient One-Shot Incremental Federated Learning
di: Zaland, Obaidullah, et al.
Pubblicazione: (2026)
di: Zaland, Obaidullah, et al.
Pubblicazione: (2026)
A Semi-Supervised Federated Learning Framework with Hierarchical Clustering Aggregation for Heterogeneous Satellite Networks
di: Liu, Zhuocheng, et al.
Pubblicazione: (2025)
di: Liu, Zhuocheng, et al.
Pubblicazione: (2025)
Clustered Federated Learning with Hierarchical Knowledge Distillation
di: Ahmad, Sabtain, et al.
Pubblicazione: (2025)
di: Ahmad, Sabtain, et al.
Pubblicazione: (2025)
FastCHGNet: Training one Universal Interatomic Potential to 1.5 Hours with 32 GPUs
di: Zhou, Yuanchang, et al.
Pubblicazione: (2024)
di: Zhou, Yuanchang, et al.
Pubblicazione: (2024)
MegaScale: Scaling Large Language Model Training to More Than 10,000 GPUs
di: Jiang, Ziheng, et al.
Pubblicazione: (2024)
di: Jiang, Ziheng, et al.
Pubblicazione: (2024)
Intelligent Orchestration of Distributed Large Foundation Model Inference at the Edge
di: Koch, Fernando, et al.
Pubblicazione: (2025)
di: Koch, Fernando, et al.
Pubblicazione: (2025)
Runtime-Orchestrated Second-Order Optimization for Scalable LLM Training
di: Lu, Yishun, et al.
Pubblicazione: (2026)
di: Lu, Yishun, et al.
Pubblicazione: (2026)
ADAPT: A Self-Calibrating Proactive Autoscaler for Container Orchestration
di: Baghel, Himanshu Singh
Pubblicazione: (2026)
di: Baghel, Himanshu Singh
Pubblicazione: (2026)
Incentivised Orchestrated Training Architecture (IOTA): A Technical Primer for Release
di: Quinque, Felix, et al.
Pubblicazione: (2025)
di: Quinque, Felix, et al.
Pubblicazione: (2025)
FlashEvolve: Accelerating Agent Self-Evolution with Asynchronous Stage Orchestration
di: Hu, Zhengding, et al.
Pubblicazione: (2026)
di: Hu, Zhengding, et al.
Pubblicazione: (2026)
DecHW: Heterogeneous Decentralized Federated Learning Exploiting Second-Order Information
di: Ahmad, Adnan, et al.
Pubblicazione: (2026)
di: Ahmad, Adnan, et al.
Pubblicazione: (2026)
Decentralized Orchestration Architecture for Fluid Computing: A Secure Distributed AI Use Case
di: Cajaraville-Aboy, Diego, et al.
Pubblicazione: (2026)
di: Cajaraville-Aboy, Diego, et al.
Pubblicazione: (2026)
SAFL: Structure-Aware Personalized Federated Learning via Client-Specific Clustering and SCSI-Guided Model Pruning
di: Li, Nan, et al.
Pubblicazione: (2025)
di: Li, Nan, et al.
Pubblicazione: (2025)
ALTO: Adaptive LoRA Tuning and Orchestration for Heterogeneous LoRA Training Workloads
di: Zuo, Jingwei, et al.
Pubblicazione: (2026)
di: Zuo, Jingwei, et al.
Pubblicazione: (2026)
Asynchronous Federated Clustering with Unknown Number of Clusters
di: Zhang, Yunfan, et al.
Pubblicazione: (2024)
di: Zhang, Yunfan, et al.
Pubblicazione: (2024)
Coral: Cost-Efficient Multi-LLM Serving over Heterogeneous Cloud GPUs
di: Mei, Yixuan, et al.
Pubblicazione: (2026)
di: Mei, Yixuan, et al.
Pubblicazione: (2026)
HDEE: Heterogeneous Domain Expert Ensemble
di: Ersoy, Oğuzhan, et al.
Pubblicazione: (2025)
di: Ersoy, Oğuzhan, et al.
Pubblicazione: (2025)
Federated K-means Clustering
di: Garst, Swier, et al.
Pubblicazione: (2023)
di: Garst, Swier, et al.
Pubblicazione: (2023)
Federated Temporal Graph Clustering
di: Zhou, Zihao, et al.
Pubblicazione: (2024)
di: Zhou, Zihao, et al.
Pubblicazione: (2024)
Heterogeneous Federated Learning with Prototype Alignment and Upscaling
di: Lee, Gyuejeong, et al.
Pubblicazione: (2025)
di: Lee, Gyuejeong, et al.
Pubblicazione: (2025)
Hypernetworks for Model-Heterogeneous Personalized Federated Learning
di: Zhang, Chen, et al.
Pubblicazione: (2025)
di: Zhang, Chen, et al.
Pubblicazione: (2025)
STHFL: Spatio-Temporal Heterogeneous Federated Learning
di: Guo, Shunxin, et al.
Pubblicazione: (2025)
di: Guo, Shunxin, et al.
Pubblicazione: (2025)
Documenti analoghi
-
HetCCL: Accelerating LLM Training with Heterogeneous GPUs
di: Kim, Heehoon, et al.
Pubblicazione: (2026) -
Tessera: Unlocking Heterogeneous GPUs through Kernel-Granularity Disaggregation
di: Hu, Tiancheng, et al.
Pubblicazione: (2026) -
HeterMoE: Efficient Training of Mixture-of-Experts Models on Heterogeneous GPUs
di: Wu, Yongji, et al.
Pubblicazione: (2025) -
AReaL-Hex: Accommodating Asynchronous RL Training over Heterogeneous GPUs
di: Yan, Ran, et al.
Pubblicazione: (2025) -
PlexRL: Cluster-Level Orchestration of Serviceized LLM Execution for RLVR
di: Zhang, Yiqi, et al.
Pubblicazione: (2026)