Salvato in:
| Autori principali: | Stubbs, Joe, Padhy, Smruti, Cardone, Richard |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2408.03349 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
GPU Cluster Scheduling for Network-Sensitive Deep Learning
di: Sharma, Aakash, et al.
Pubblicazione: (2024)
di: Sharma, Aakash, et al.
Pubblicazione: (2024)
Is Intelligence the Right Direction in New OS Scheduling for Multiple Resources in Cloud Environments?
di: Dou, Xinglei, et al.
Pubblicazione: (2025)
di: Dou, Xinglei, et al.
Pubblicazione: (2025)
Towards Universal Performance Modeling for Machine Learning Training on Multi-GPU Platforms
di: Lin, Zhongyi, et al.
Pubblicazione: (2024)
di: Lin, Zhongyi, et al.
Pubblicazione: (2024)
Prompt-Aware Scheduling for Low-Latency LLM Serving
di: Tao, Yiheng, et al.
Pubblicazione: (2025)
di: Tao, Yiheng, et al.
Pubblicazione: (2025)
Agentic Auto-Scheduling: An Experimental Study of LLM-Guided Loop Optimization
di: Merouani, Massinissa, et al.
Pubblicazione: (2025)
di: Merouani, Massinissa, et al.
Pubblicazione: (2025)
A Comparative Study of OpenMP Scheduling Algorithm Selection Strategies
di: Korndörfer, Jonas H. Müller, et al.
Pubblicazione: (2025)
di: Korndörfer, Jonas H. Müller, et al.
Pubblicazione: (2025)
KVDirect: Distributed Disaggregated LLM Inference
di: Chen, Shiyang, et al.
Pubblicazione: (2024)
di: Chen, Shiyang, et al.
Pubblicazione: (2024)
KVPR: Efficient LLM Inference with I/O-Aware KV Cache Partial Recomputation
di: Jiang, Chaoyi, et al.
Pubblicazione: (2024)
di: Jiang, Chaoyi, et al.
Pubblicazione: (2024)
When Less is More: Achieving Faster Convergence in Distributed Edge Machine Learning
di: Basani, Advik Raj, et al.
Pubblicazione: (2024)
di: Basani, Advik Raj, et al.
Pubblicazione: (2024)
iSpLib: A Library for Accelerating Graph Neural Networks using Auto-tuned Sparse Operations
di: Anik, Md Saidul Hoque, et al.
Pubblicazione: (2024)
di: Anik, Md Saidul Hoque, et al.
Pubblicazione: (2024)
cedar: Optimized and Unified Machine Learning Input Data Pipelines
di: Zhao, Mark, et al.
Pubblicazione: (2024)
di: Zhao, Mark, et al.
Pubblicazione: (2024)
MIREncoder: Multi-modal IR-based Pretrained Embeddings for Performance Optimizations
di: Dutta, Akash, et al.
Pubblicazione: (2024)
di: Dutta, Akash, et al.
Pubblicazione: (2024)
AutoChunk: Automated Activation Chunk for Memory-Efficient Long Sequence Inference
di: Zhao, Xuanlei, et al.
Pubblicazione: (2024)
di: Zhao, Xuanlei, et al.
Pubblicazione: (2024)
MassiveGNN: Efficient Training via Prefetching for Massively Connected Distributed Graphs
di: Sarkar, Aishwarya, et al.
Pubblicazione: (2024)
di: Sarkar, Aishwarya, et al.
Pubblicazione: (2024)
Less is More: Optimizing Function Calling for LLM Execution on Edge Devices
di: Paramanayakam, Varatheepan, et al.
Pubblicazione: (2024)
di: Paramanayakam, Varatheepan, et al.
Pubblicazione: (2024)
BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures
di: Hu, Xiannan, et al.
Pubblicazione: (2025)
di: Hu, Xiannan, et al.
Pubblicazione: (2025)
Fake Runs, Real Fixes -- Analyzing xPU Performance Through Simulation
di: Zarkadas, Ioannis, et al.
Pubblicazione: (2025)
di: Zarkadas, Ioannis, et al.
Pubblicazione: (2025)
InkStream: Real-time GNN Inference on Streaming Graphs via Incremental Update
di: Wu, Dan, et al.
Pubblicazione: (2023)
di: Wu, Dan, et al.
Pubblicazione: (2023)
Vectorized FlashAttention with Low-cost Exponential Computation in RISC-V Vector Processors
di: Titopoulos, Vasileios, et al.
Pubblicazione: (2025)
di: Titopoulos, Vasileios, et al.
Pubblicazione: (2025)
TaxBreak: Unmasking the Hidden Costs of LLM Inference Through Overhead Decomposition
di: Vellaisamy, Prabhu, et al.
Pubblicazione: (2026)
di: Vellaisamy, Prabhu, et al.
Pubblicazione: (2026)
DeepCQ: General-Purpose Deep-Surrogate Framework for Lossy Compression Quality Prediction
di: Mumenin, Khondoker Mirazul, et al.
Pubblicazione: (2025)
di: Mumenin, Khondoker Mirazul, et al.
Pubblicazione: (2025)
xMem: A CPU-Based Approach for Accurate Estimation of GPU Memory in Deep Learning Training Workloads
di: Shi, Jiabo, et al.
Pubblicazione: (2025)
di: Shi, Jiabo, et al.
Pubblicazione: (2025)
IPA: Inference Pipeline Adaptation to Achieve High Accuracy and Cost-Efficiency
di: Ghafouri, Saeid, et al.
Pubblicazione: (2023)
di: Ghafouri, Saeid, et al.
Pubblicazione: (2023)
CoFormer: Collaborating with Heterogeneous Edge Devices for Scalable Transformer Inference
di: Xu, Guanyu, et al.
Pubblicazione: (2025)
di: Xu, Guanyu, et al.
Pubblicazione: (2025)
Accelerating Mobile Inference through Fine-Grained CPU-GPU Co-Execution
di: Li, Zhuojin, et al.
Pubblicazione: (2025)
di: Li, Zhuojin, et al.
Pubblicazione: (2025)
Multi-DNN Inference of Sparse Models on Edge SoCs
di: Luo, Jiawei, et al.
Pubblicazione: (2026)
di: Luo, Jiawei, et al.
Pubblicazione: (2026)
MLCommons Chakra: Advancing Performance Benchmarking and Co-design using Standardized Execution Traces
di: Sridharan, Srinivas, et al.
Pubblicazione: (2026)
di: Sridharan, Srinivas, et al.
Pubblicazione: (2026)
AutoSP: Unlocking Long-Context LLM Training Via Compiler-Based Sequence Parallelism
di: Gupta, Ahan, et al.
Pubblicazione: (2026)
di: Gupta, Ahan, et al.
Pubblicazione: (2026)
ReLATE: Learning Efficient Sparse Encoding for High-Performance Tensor Decomposition
di: Helal, Ahmed E., et al.
Pubblicazione: (2025)
di: Helal, Ahmed E., et al.
Pubblicazione: (2025)
Execution time budget assignment for mixed criticality systems
di: Khelassi, Mohamed Amine, et al.
Pubblicazione: (2023)
di: Khelassi, Mohamed Amine, et al.
Pubblicazione: (2023)
Ecomap: Sustainability-Driven Optimization of Multi-Tenant DNN Execution on Edge Servers
di: Paramanayakam, Varatheepan, et al.
Pubblicazione: (2025)
di: Paramanayakam, Varatheepan, et al.
Pubblicazione: (2025)
Distributed Matrix-Based Sampling for Graph Neural Network Training
di: Tripathy, Alok, et al.
Pubblicazione: (2023)
di: Tripathy, Alok, et al.
Pubblicazione: (2023)
CloudFormer: An Attention-based Performance Prediction for Public Clouds with Unknown Workload
di: Shahbazinia, Amirhossein, et al.
Pubblicazione: (2025)
di: Shahbazinia, Amirhossein, et al.
Pubblicazione: (2025)
Glinthawk: A Two-Tiered Architecture for Offline LLM Inference
di: Hamadanian, Pouya, et al.
Pubblicazione: (2025)
di: Hamadanian, Pouya, et al.
Pubblicazione: (2025)
CARMA: Collocation-Aware Resource Manager
di: Yousefzadeh-Asl-Miandoab, Ehsan, et al.
Pubblicazione: (2025)
di: Yousefzadeh-Asl-Miandoab, Ehsan, et al.
Pubblicazione: (2025)
Ariel-ML: Computing Parallelization with Embedded Rust for Neural Networks on Heterogeneous Multi-core Microcontrollers
di: Huang, Zhaolan, et al.
Pubblicazione: (2025)
di: Huang, Zhaolan, et al.
Pubblicazione: (2025)
You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models
di: Ding, Shiwei, et al.
Pubblicazione: (2025)
di: Ding, Shiwei, et al.
Pubblicazione: (2025)
Tuning the Tuner: Introducing Hyperparameter Optimization for Auto-Tuning
di: Willemsen, Floris-Jan, et al.
Pubblicazione: (2025)
di: Willemsen, Floris-Jan, et al.
Pubblicazione: (2025)
A Practical Two-Stage Framework for GPU Resource and Power Prediction in Heterogeneous HPC Systems
di: Oztop, Beste, et al.
Pubblicazione: (2026)
di: Oztop, Beste, et al.
Pubblicazione: (2026)
Characterizing WebGPU Dispatch Overhead for LLM Inference Across Four GPU Vendors, Three Backends, and Three Browsers
di: Maczan, Jędrzej
Pubblicazione: (2026)
di: Maczan, Jędrzej
Pubblicazione: (2026)
Documenti analoghi
-
GPU Cluster Scheduling for Network-Sensitive Deep Learning
di: Sharma, Aakash, et al.
Pubblicazione: (2024) -
Is Intelligence the Right Direction in New OS Scheduling for Multiple Resources in Cloud Environments?
di: Dou, Xinglei, et al.
Pubblicazione: (2025) -
Towards Universal Performance Modeling for Machine Learning Training on Multi-GPU Platforms
di: Lin, Zhongyi, et al.
Pubblicazione: (2024) -
Prompt-Aware Scheduling for Low-Latency LLM Serving
di: Tao, Yiheng, et al.
Pubblicazione: (2025) -
Agentic Auto-Scheduling: An Experimental Study of LLM-Guided Loop Optimization
di: Merouani, Massinissa, et al.
Pubblicazione: (2025)