Tesserae: Scalable Placement Policies for Deep Learning Workloads
Fuente:
arXiv
Saved in:
| Main Authors: | Bian, Song, Agarwal, Saurabh, Mahmood, Md. Tareq, Venkataraman, Shivaram |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SYMPHONY: Improving Memory Management for LLM Inference Workloads
by: Agarwal, Saurabh, et al.
Published: (2024)
by: Agarwal, Saurabh, et al.
Published: (2024)
PAL: A Variability-Aware Policy for Scheduling ML Workloads in GPU Clusters
by: Jain, Rutwik, et al.
Published: (2024)
by: Jain, Rutwik, et al.
Published: (2024)
Minos: Systematically Classifying Performance and Power Characteristics of GPU Workloads on HPC Clusters
by: Jain, Rutwik, et al.
Published: (2026)
by: Jain, Rutwik, et al.
Published: (2026)
Eva: Cost-Efficient Cloud-Based Cluster Scheduling
by: Chang, Tzu-Tao, et al.
Published: (2025)
by: Chang, Tzu-Tao, et al.
Published: (2025)
Resource Allocation and Workload Scheduling for Large-Scale Distributed Deep Learning: A Survey
by: Liang, Feng, et al.
Published: (2024)
by: Liang, Feng, et al.
Published: (2024)
LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models
by: Chang, Tzu-Tao, et al.
Published: (2025)
by: Chang, Tzu-Tao, et al.
Published: (2025)
Placement Semantics for Distributed Deep Learning: A Systematic Framework for Analyzing Parallelism Strategies
by: Mehta, Deep Pankajbhai
Published: (2026)
by: Mehta, Deep Pankajbhai
Published: (2026)
Quality Scalable Quantization Methodology for Deep Learning on Edge
by: Khaliq, Salman Abdul, et al.
Published: (2024)
by: Khaliq, Salman Abdul, et al.
Published: (2024)
Duration-Informed Workload Scheduler
by: Loreti, Daniela, et al.
Published: (2026)
by: Loreti, Daniela, et al.
Published: (2026)
PGT-I: Scaling Spatiotemporal GNNs with Memory-Efficient Distributed Training
by: Ockerman, Seth, et al.
Published: (2025)
by: Ockerman, Seth, et al.
Published: (2025)
Workload Schedulers -- Genesis, Algorithms and Differences
by: Sliwko, Leszek, et al.
Published: (2025)
by: Sliwko, Leszek, et al.
Published: (2025)
Towards Carbon-Aware Container Orchestration: Predicting Workload Energy Consumption with Federated Learning
by: Saad, Zainab, et al.
Published: (2025)
by: Saad, Zainab, et al.
Published: (2025)
Topology-aware Preemptive Scheduling for Co-located LLM Workloads
by: Zhang, Ping, et al.
Published: (2024)
by: Zhang, Ping, et al.
Published: (2024)
Scalable Cloud-Native Architectures for Intelligent PMU Data Processing
by: Chockalingam, Nachiappan, et al.
Published: (2025)
by: Chockalingam, Nachiappan, et al.
Published: (2025)
HiveMind: OS-Inspired Scheduling for Concurrent LLM Agent Workloads
by: Agyemang, Justice Owusu, et al.
Published: (2026)
by: Agyemang, Justice Owusu, et al.
Published: (2026)
Hubs and Spokes Learning: Efficient and Scalable Collaborative Machine Learning
by: Sharma, Atul, et al.
Published: (2025)
by: Sharma, Atul, et al.
Published: (2025)
Hybrid Heterogeneous Clusters Can Lower the Energy Consumption of LLM Inference Workloads
by: Wilkins, Grant, et al.
Published: (2024)
by: Wilkins, Grant, et al.
Published: (2024)
Optimal Workload Placement on Multi-Instance GPUs
by: Turkkan, Bekir, et al.
Published: (2024)
by: Turkkan, Bekir, et al.
Published: (2024)
Mixture-of-Schedulers: An Adaptive Scheduling Agent as a Learned Router for Expert Policies
by: Wang, Xinbo, et al.
Published: (2025)
by: Wang, Xinbo, et al.
Published: (2025)
Quantifying Energy and Cost Benefits of Hybrid Edge Cloud: Analysis of Traditional and Agentic Workloads
by: Alamouti, Siavash
Published: (2025)
by: Alamouti, Siavash
Published: (2025)
Rethinking Inference Placement for Deep Learning across Edge and Cloud Platforms: A Multi-Objective Optimization Perspective and Future Directions
by: Zhang, Zongshun, et al.
Published: (2025)
by: Zhang, Zongshun, et al.
Published: (2025)
ReCoVer: Resilient LLM Pre-Training System via Fault-Tolerant Collective and Versatile Workload
by: Liu, Ziyue, et al.
Published: (2026)
by: Liu, Ziyue, et al.
Published: (2026)
On-demand Cold Start Frequency Reduction with Off-Policy Reinforcement Learning in Serverless Computing
by: Agarwal, Siddharth, et al.
Published: (2023)
by: Agarwal, Siddharth, et al.
Published: (2023)
Enabling Secure and Ephemeral AI Workloads in Data Mesh Environments
by: Patel, Chinkit, et al.
Published: (2025)
by: Patel, Chinkit, et al.
Published: (2025)
Privacy-Preserving Federated Learning: Integrating Zero-Knowledge Proofs in Scalable Distributed Architectures
by: Gupta, Divya
Published: (2026)
by: Gupta, Divya
Published: (2026)
Safactory: A Scalable Agentic Infrastructure for Training Trustworthy Autonomous Intelligence
by: Chen, Xinquan, et al.
Published: (2026)
by: Chen, Xinquan, et al.
Published: (2026)
A Deep Recurrent-Reinforcement Learning Method for Intelligent AutoScaling of Serverless Functions
by: Agarwal, Siddharth, et al.
Published: (2023)
by: Agarwal, Siddharth, et al.
Published: (2023)
The Big Send-off: Scalable and Performant Collectives for Deep Learning
by: Singh, Siddharth, et al.
Published: (2025)
by: Singh, Siddharth, et al.
Published: (2025)
StreamServe: Adaptive Speculative Flows for Low-Latency Disaggregated LLM Serving
by: Kumar, Satyam, et al.
Published: (2026)
by: Kumar, Satyam, et al.
Published: (2026)
Game-Theoretic Deep Reinforcement Learning to Minimize Carbon Emissions and Energy Costs for AI Inference Workloads in Geo-Distributed Data Centers
by: Hogade, Ninad, et al.
Published: (2024)
by: Hogade, Ninad, et al.
Published: (2024)
An Advanced Reinforcement Learning Framework for Online Scheduling of Deferrable Workloads in Cloud Computing
by: Dong, Hang, et al.
Published: (2024)
by: Dong, Hang, et al.
Published: (2024)
High-Dimensional Data Processing: Benchmarking Machine Learning and Deep Learning Architectures in Local and Distributed Environments
by: Rodriguez, Julian, et al.
Published: (2025)
by: Rodriguez, Julian, et al.
Published: (2025)
Accurate GPU Memory Prediction for Deep Learning Jobs through Dynamic Analysis
by: Shi, Jiabo, et al.
Published: (2025)
by: Shi, Jiabo, et al.
Published: (2025)
Verify Distributed Deep Learning Model Implementation Refinement with Iterative Relation Inference
by: Wang, Zhanghan, et al.
Published: (2025)
by: Wang, Zhanghan, et al.
Published: (2025)
Deep Reinforcement Learning for Fault-Adaptive Routing in Eisenstein-Jacobi Interconnection Topologies
by: Charrwi, Mohammad Walid, et al.
Published: (2026)
by: Charrwi, Mohammad Walid, et al.
Published: (2026)
Communication-Efficient Large-Scale Distributed Deep Learning: A Comprehensive Survey
by: Liang, Feng, et al.
Published: (2024)
by: Liang, Feng, et al.
Published: (2024)
Hierarchical Autoscaling for Large Language Model Serving with Chiron
by: Patke, Archit, et al.
Published: (2025)
by: Patke, Archit, et al.
Published: (2025)
Hybrid Learning and Optimization-Based Dynamic Scheduling for DL Workloads on Heterogeneous GPU Clusters
by: Dongare, Shruti, et al.
Published: (2025)
by: Dongare, Shruti, et al.
Published: (2025)
TensorHub: Scalable and Elastic Weight Transfer for LLM RL Training
by: Ye, Chenhao, et al.
Published: (2026)
by: Ye, Chenhao, et al.
Published: (2026)
Towards Scalable GPU-Accelerated SNN Training via Temporal Fusion
by: Li, Yanchen, et al.
Published: (2024)
by: Li, Yanchen, et al.
Published: (2024)
Similar Items
-
SYMPHONY: Improving Memory Management for LLM Inference Workloads
by: Agarwal, Saurabh, et al.
Published: (2024) -
PAL: A Variability-Aware Policy for Scheduling ML Workloads in GPU Clusters
by: Jain, Rutwik, et al.
Published: (2024) -
Minos: Systematically Classifying Performance and Power Characteristics of GPU Workloads on HPC Clusters
by: Jain, Rutwik, et al.
Published: (2026) -
Eva: Cost-Efficient Cloud-Based Cluster Scheduling
by: Chang, Tzu-Tao, et al.
Published: (2025) -
Resource Allocation and Workload Scheduling for Large-Scale Distributed Deep Learning: A Survey
by: Liang, Feng, et al.
Published: (2024)