Saved in:
| Main Authors: | Sun, Wenbo, Hai, Rihan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2502.01985 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SiloFuse: Cross-silo Synthetic Data Generation with Latent Tabular Diffusion Models
by: Shankar, Aditya, et al.
Published: (2024)
by: Shankar, Aditya, et al.
Published: (2024)
Towards Universal Performance Modeling for Machine Learning Training on Multi-GPU Platforms
by: Lin, Zhongyi, et al.
Published: (2024)
by: Lin, Zhongyi, et al.
Published: (2024)
Magneton: Optimizing Energy Efficiency of ML Systems via Differential Energy Debugging
by: Pan, Yi, et al.
Published: (2025)
by: Pan, Yi, et al.
Published: (2025)
A Distributed Framework for Causal Modeling of Performance Variability in GPU Traces
by: Lahiry, Ankur, et al.
Published: (2025)
by: Lahiry, Ankur, et al.
Published: (2025)
Taming the Memory Beast: Strategies for Reliable ML Training on Kubernetes
by: Ray, Jaideep
Published: (2024)
by: Ray, Jaideep
Published: (2024)
Share Secrets for Privacy: Confidential Forecasting with Vertical Federated Learning
by: Shankar, Aditya, et al.
Published: (2024)
by: Shankar, Aditya, et al.
Published: (2024)
Semantic-Aware Scheduling for GPU Clusters with Large Language Models
by: Wang, Zerui, et al.
Published: (2025)
by: Wang, Zerui, et al.
Published: (2025)
Incentive-Compatible Federated Learning with Stackelberg Game Modeling
by: Javaherian, Simin, et al.
Published: (2025)
by: Javaherian, Simin, et al.
Published: (2025)
Maya: Optimizing Deep Learning Training Workloads using GPU Runtime Emulation
by: Yarlagadda, Srihas, et al.
Published: (2025)
by: Yarlagadda, Srihas, et al.
Published: (2025)
GPU Memory Prediction for Multimodal Model Training
by: Jeong, Jinwoo, et al.
Published: (2025)
by: Jeong, Jinwoo, et al.
Published: (2025)
Nonuniform-Tensor-Parallelism: Mitigating GPU failure impact for Scaled-up LLM Training
by: Arfeen, Daiyaan, et al.
Published: (2025)
by: Arfeen, Daiyaan, et al.
Published: (2025)
LSM-GNN: Large-scale Storage-based Multi-GPU GNN Training by Optimizing Data Transfer Scheme
by: Park, Jeongmin Brian, et al.
Published: (2024)
by: Park, Jeongmin Brian, et al.
Published: (2024)
SERFLOW: A Cross-Service Cost Optimization Framework for SLO-Aware Dynamic ML Inference
by: Zhang, Zongshun, et al.
Published: (2025)
by: Zhang, Zongshun, et al.
Published: (2025)
WarmServe: Enabling One-for-Many GPU Prewarming for Multi-LLM Serving
by: Lou, Chiheng, et al.
Published: (2025)
by: Lou, Chiheng, et al.
Published: (2025)
Mélange: Cost Efficient Large Language Model Serving by Exploiting GPU Heterogeneity
by: Griggs, Tyler, et al.
Published: (2024)
by: Griggs, Tyler, et al.
Published: (2024)
HybriMoE: Hybrid CPU-GPU Scheduling and Cache Management for Efficient MoE Inference
by: Zhong, Shuzhang, et al.
Published: (2025)
by: Zhong, Shuzhang, et al.
Published: (2025)
Single-GPU GNN Systems: Traps and Pitfalls
by: Gong, Yidong, et al.
Published: (2024)
by: Gong, Yidong, et al.
Published: (2024)
BOOST: BOttleneck-Optimized Scalable Training Framework for Low-Rank Large Language Models
by: Wang, Zhengyang, et al.
Published: (2025)
by: Wang, Zhengyang, et al.
Published: (2025)
An All-Reduce Compatible Top-K Compressor for Communication-Efficient Distributed Learning
by: Chen, Chuyan, et al.
Published: (2025)
by: Chen, Chuyan, et al.
Published: (2025)
FastGL: A GPU-Efficient Framework for Accelerating Sampling-Based GNN Training at Large Scale
by: Zhu, Zeyu, et al.
Published: (2024)
by: Zhu, Zeyu, et al.
Published: (2024)
A Robust Power Model Training Framework for Cloud Native Runtime Energy Metric Exporter
by: Choochotkaew, Sunyanan, et al.
Published: (2024)
by: Choochotkaew, Sunyanan, et al.
Published: (2024)
LLMPerf: GPU Performance Modeling meets Large Language Models
by: Nguyen, Khoi N. M., et al.
Published: (2025)
by: Nguyen, Khoi N. M., et al.
Published: (2025)
Parm: Efficient Training of Large Sparsely-Activated Models with Dedicated Schedules
by: Pan, Xinglin, et al.
Published: (2024)
by: Pan, Xinglin, et al.
Published: (2024)
Iris: First-Class Multi-GPU Programming Experience in Triton
by: Awad, Muhammad, et al.
Published: (2025)
by: Awad, Muhammad, et al.
Published: (2025)
KernelFoundry: Hardware-aware evolutionary GPU kernel optimization
by: Wiedemann, Nina, et al.
Published: (2026)
by: Wiedemann, Nina, et al.
Published: (2026)
GNNBENCH: Fair and Productive Benchmarking for Single-GPU GNN System
by: Gong, Yidong, et al.
Published: (2024)
by: Gong, Yidong, et al.
Published: (2024)
Modalities, a PyTorch-native Framework For Large-scale LLM Training and Research
by: Lübbering, Max, et al.
Published: (2026)
by: Lübbering, Max, et al.
Published: (2026)
A Bring-Your-Own-Model Approach for ML-Driven Storage Placement in Warehouse-Scale Computers
by: Yang, Chenxi, et al.
Published: (2025)
by: Yang, Chenxi, et al.
Published: (2025)
Keras Sig: Efficient Path Signature Computation on GPU in Keras 3
by: Genet, Rémi, et al.
Published: (2025)
by: Genet, Rémi, et al.
Published: (2025)
ParallelKittens: Systematic and Practical Simplification of Multi-GPU AI Kernels
by: Sul, Stuart H., et al.
Published: (2025)
by: Sul, Stuart H., et al.
Published: (2025)
Nexus:Proactive Intra-GPU Disaggregation of Prefill and Decode in LLM Serving
by: Shi, Xiaoxiang, et al.
Published: (2025)
by: Shi, Xiaoxiang, et al.
Published: (2025)
Challenging GPU Dominance: When CPUs Outperform for On-Device LLM Inference
by: Zhang, Haolin, et al.
Published: (2025)
by: Zhang, Haolin, et al.
Published: (2025)
Efficient Remote KV Cache Reuse with GPU-native Video Codec
by: Mi, Liang, et al.
Published: (2026)
by: Mi, Liang, et al.
Published: (2026)
Revati: Transparent GPU-Free Time-Warp Emulation for LLM Serving
by: Agrawal, Amey, et al.
Published: (2026)
by: Agrawal, Amey, et al.
Published: (2026)
Parallel Track Transformers: Enabling Fast GPU Inference with Reduced Synchronization
by: Wang, Chong, et al.
Published: (2026)
by: Wang, Chong, et al.
Published: (2026)
xMem: A CPU-Based Approach for Accurate Estimation of GPU Memory in Deep Learning Training Workloads
by: Shi, Jiabo, et al.
Published: (2025)
by: Shi, Jiabo, et al.
Published: (2025)
DORA: A Scalable Asynchronous Reinforcement Learning System for Language Model Training
by: Hu, Tianhao, et al.
Published: (2026)
by: Hu, Tianhao, et al.
Published: (2026)
RollPacker: Mitigating Long-Tail Rollouts for Fast, Synchronous RL Post-Training
by: Gao, Wei, et al.
Published: (2025)
by: Gao, Wei, et al.
Published: (2025)
Adacc: An Adaptive Framework Unifying Compression and Activation Recomputation for LLM Training
by: Chen, Ping, et al.
Published: (2025)
by: Chen, Ping, et al.
Published: (2025)
A Practical Two-Stage Framework for GPU Resource and Power Prediction in Heterogeneous HPC Systems
by: Oztop, Beste, et al.
Published: (2026)
by: Oztop, Beste, et al.
Published: (2026)
Similar Items
-
SiloFuse: Cross-silo Synthetic Data Generation with Latent Tabular Diffusion Models
by: Shankar, Aditya, et al.
Published: (2024) -
Towards Universal Performance Modeling for Machine Learning Training on Multi-GPU Platforms
by: Lin, Zhongyi, et al.
Published: (2024) -
Magneton: Optimizing Energy Efficiency of ML Systems via Differential Energy Debugging
by: Pan, Yi, et al.
Published: (2025) -
A Distributed Framework for Causal Modeling of Performance Variability in GPU Traces
by: Lahiry, Ankur, et al.
Published: (2025) -
Taming the Memory Beast: Strategies for Reliable ML Training on Kubernetes
by: Ray, Jaideep
Published: (2024)