Evaluating Large Language Models for Workload Mapping and Scheduling in Heterogeneous HPC Systems
Fuente:
arXiv
Saved in:
| Main Authors: | Sharma, Aasish Kumar, Kunkel, Julian |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
An Empirical Evaluation of Quantum-Inspired QUBO Methods for Heterogeneous HPC Workflow Mapping and Scheduling
by: Sharma, Aasish Kumar, et al.
Published: (2026)
by: Sharma, Aasish Kumar, et al.
Published: (2026)
ConfigSpec: Profiling-Based Configuration Selection for Distributed Edge--Cloud Speculative LLM Serving
by: Li, Xiangchen, et al.
Published: (2026)
by: Li, Xiangchen, et al.
Published: (2026)
WISP: Waste- and Interference-Suppressed Distributed Speculative LLM Serving at the Edge via Dynamic Drafting and SLO-Aware Batching
by: Li, Xiangchen, et al.
Published: (2026)
by: Li, Xiangchen, et al.
Published: (2026)
Ontological Knowledge Blocks: Executable Compliance and Profile-Based Validation for Trustworthy AI Systems
by: Sharma, Aasish Kumar, et al.
Published: (2026)
by: Sharma, Aasish Kumar, et al.
Published: (2026)
Scalability Optimization in Cloud-Based AI Inference Services: Strategies for Real-Time Load Balancing and Automated Scaling
by: Jin, Yihong, et al.
Published: (2025)
by: Jin, Yihong, et al.
Published: (2025)
Experimentally Evaluating the Resource Efficiency of Big Data Autoscaling
by: Will, Jonathan, et al.
Published: (2025)
by: Will, Jonathan, et al.
Published: (2025)
Deadline-Aware Joint Task Scheduling and Offloading in Mobile Edge Computing Systems
by: Nguyen, Ngoc Hung, et al.
Published: (2025)
by: Nguyen, Ngoc Hung, et al.
Published: (2025)
A Review of Tools and Techniques for Optimization of Workload Mapping and Scheduling in Heterogeneous HPC System
by: Sharma, Aasish Kumar, et al.
Published: (2025)
by: Sharma, Aasish Kumar, et al.
Published: (2025)
Efficient Construction of Large Search Spaces for Auto-Tuning
by: Willemsen, Floris-Jan, et al.
Published: (2025)
by: Willemsen, Floris-Jan, et al.
Published: (2025)
Spark-LLM-Eval: A Distributed Framework for Statistically Rigorous Large Language Model Evaluation
by: Mitra, Subhadip
Published: (2026)
by: Mitra, Subhadip
Published: (2026)
Scalable overset computation between a forest-of-octrees- and an arbitrary distributed parallel mesh
by: Brandt, Hannes, et al.
Published: (2026)
by: Brandt, Hannes, et al.
Published: (2026)
Learning Interpretable Scheduling Algorithms for Data Processing Clusters
by: Hu, Zhibo, et al.
Published: (2024)
by: Hu, Zhibo, et al.
Published: (2024)
FedMon: Federated eBPF Monitoring for Distributed Anomaly Detection in Multi-Cluster Cloud Environments
by: Zehra, Sehar, et al.
Published: (2025)
by: Zehra, Sehar, et al.
Published: (2025)
Deploy, Calibrate, Monitor, Heal -- No Human Required: An Autonomous AI SRE Agent for Elasticsearch
by: Mukkolakkal, Muhamed Ramees Cheriya
Published: (2026)
by: Mukkolakkal, Muhamed Ramees Cheriya
Published: (2026)
DPDPU: Data Processing with DPUs
by: Hu, Jiasheng, et al.
Published: (2024)
by: Hu, Jiasheng, et al.
Published: (2024)
Efficiently Scheduling Parallel DAG Tasks on Identical Multiprocessors
by: Lendve, Shardul, et al.
Published: (2024)
by: Lendve, Shardul, et al.
Published: (2024)
FLEdge: Benchmarking Federated Machine Learning Applications in Edge Computing Systems
by: Woisetschläger, Herbert, et al.
Published: (2023)
by: Woisetschläger, Herbert, et al.
Published: (2023)
GraphBit: A Graph-based Agentic Framework for Non-Linear Agent Orchestration
by: Sarker, Yeahia, et al.
Published: (2026)
by: Sarker, Yeahia, et al.
Published: (2026)
Benchmarking Federated Learning for Throughput Prediction in 5G Live Streaming Applications
by: Dutta, Yuvraj, et al.
Published: (2025)
by: Dutta, Yuvraj, et al.
Published: (2025)
Scalable Engine and the Performance of Different LLM Models in a SLURM based HPC architecture
by: Luiz, Anderson de Lima, et al.
Published: (2025)
by: Luiz, Anderson de Lima, et al.
Published: (2025)
Heuristic Search Space Partitioning for Low-Latency Multi-Tenant Cloud Queries
by: Pathak, Prashant Kumar, et al.
Published: (2026)
by: Pathak, Prashant Kumar, et al.
Published: (2026)
AAFLOW: Scalable Patterns for Agentic AI Workflows
by: Sarker, Arup Kumar, et al.
Published: (2026)
by: Sarker, Arup Kumar, et al.
Published: (2026)
Design and Implementation of an Analysis Pipeline for Heterogeneous Data
by: Sarker, Arup Kumar, et al.
Published: (2024)
by: Sarker, Arup Kumar, et al.
Published: (2024)
Kant: An Efficient Unified Scheduling System for Large-Scale AI Clusters
by: Zeng, Lingling, et al.
Published: (2025)
by: Zeng, Lingling, et al.
Published: (2025)
Federated Fine-Tuning of LLMs on the Very Edge: The Good, the Bad, the Ugly
by: Woisetschläger, Herbert, et al.
Published: (2023)
by: Woisetschläger, Herbert, et al.
Published: (2023)
Service Discovery-Based Hybrid Network Middleware for Efficient Communication in Distributed Robotic Systems
by: Sang, Shiyao, et al.
Published: (2025)
by: Sang, Shiyao, et al.
Published: (2025)
Vertical Federated Image Segmentation
by: Mandal, Paul K., et al.
Published: (2024)
by: Mandal, Paul K., et al.
Published: (2024)
Horizontal Federated Computer Vision
by: Mandal, Paul K., et al.
Published: (2023)
by: Mandal, Paul K., et al.
Published: (2023)
CooperLLM: Cloud-Edge-End Cooperative Federated Fine-tuning for LLMs via ZOO-based Gradient Correction
by: Sun, He, et al.
Published: (2026)
by: Sun, He, et al.
Published: (2026)
Training LLMs on HPC Systems: Best Practices from the OpenGPT-X Project
by: Penke, Carolin, et al.
Published: (2025)
by: Penke, Carolin, et al.
Published: (2025)
Operational Memory Architecture for Kubernetes:Preserving Causal Context Across the Evidence Horizon
by: Khan, Shamsher
Published: (2026)
by: Khan, Shamsher
Published: (2026)
POD-Attention: Unlocking Full Prefill-Decode Overlap for Faster LLM Inference
by: Kamath, Aditya K, et al.
Published: (2024)
by: Kamath, Aditya K, et al.
Published: (2024)
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended
by: Kamath, Aditya K, et al.
Published: (2026)
by: Kamath, Aditya K, et al.
Published: (2026)
Trident: Adaptive Scheduling for Heterogeneous Multimodal Data Pipelines
by: Pan, Ding, et al.
Published: (2026)
by: Pan, Ding, et al.
Published: (2026)
Flex-MIG: Enabling Distributed Execution on MIG
by: Kim, Myeongsu, et al.
Published: (2025)
by: Kim, Myeongsu, et al.
Published: (2025)
nvidia-pcm: A D-Bus-Driven Platform Configuration Manager for OpenBMC Environments
by: Singh, Harinder
Published: (2026)
by: Singh, Harinder
Published: (2026)
ATTNChecker: Highly-Optimized Fault Tolerant Attention for Large Language Model Training
by: Liang, Yuhang, et al.
Published: (2024)
by: Liang, Yuhang, et al.
Published: (2024)
Rotary GPU: Exploring Local Execution Paths for Large Mixture-of-Experts Models Under Limited GPU Memory
by: Jo, Myeong Jun
Published: (2026)
by: Jo, Myeong Jun
Published: (2026)
Cost-Aware Logging: Measuring the Financial Impact of Excessive Log Retention in Small-Scale Cloud Deployments
by: Putra, Jody Almaida
Published: (2026)
by: Putra, Jody Almaida
Published: (2026)
push0: Scalable and Fault-Tolerant Orchestration for Zero-Knowledge Proof Generation
by: Ahmadvand, Mohsen, et al.
Published: (2026)
by: Ahmadvand, Mohsen, et al.
Published: (2026)
Similar Items
-
An Empirical Evaluation of Quantum-Inspired QUBO Methods for Heterogeneous HPC Workflow Mapping and Scheduling
by: Sharma, Aasish Kumar, et al.
Published: (2026) -
ConfigSpec: Profiling-Based Configuration Selection for Distributed Edge--Cloud Speculative LLM Serving
by: Li, Xiangchen, et al.
Published: (2026) -
WISP: Waste- and Interference-Suppressed Distributed Speculative LLM Serving at the Edge via Dynamic Drafting and SLO-Aware Batching
by: Li, Xiangchen, et al.
Published: (2026) -
Ontological Knowledge Blocks: Executable Compliance and Profile-Based Validation for Trustworthy AI Systems
by: Sharma, Aasish Kumar, et al.
Published: (2026) -
Scalability Optimization in Cloud-Based AI Inference Services: Strategies for Real-Time Load Balancing and Automated Scaling
by: Jin, Yihong, et al.
Published: (2025)