MARLIN: Multi-Agent Game-Theoretic Reinforcement Learning for Sustainable LLM Inference in Cloud Datacenters
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Moore, H., Qi, S., Milojicic, D., Bash, C., Pasricha, S. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Sustainable Carbon-Aware and Water-Efficient LLM Scheduling in Geo-Distributed Cloud Datacenters
von: Moore, Hayden, et al.
Veröffentlicht: (2025)
von: Moore, Hayden, et al.
Veröffentlicht: (2025)
Sustainable Graph Analytics Workload Scheduling with Evolutionary Reinforcement Learning in Edge-Cloud Systems
von: Ramicetty, P., et al.
Veröffentlicht: (2026)
von: Ramicetty, P., et al.
Veröffentlicht: (2026)
A Framework for SLO, Carbon, and Wastewater-Aware Sustainable FaaS Cloud Platform Management
von: Qi, Sirui, et al.
Veröffentlicht: (2024)
von: Qi, Sirui, et al.
Veröffentlicht: (2024)
CASA: A Framework for SLO and Carbon-Aware Autoscaling and Scheduling in Serverless Cloud Computing
von: Qi, S., et al.
Veröffentlicht: (2024)
von: Qi, S., et al.
Veröffentlicht: (2024)
Game-Theoretic Deep Reinforcement Learning to Minimize Carbon Emissions and Energy Costs for AI Inference Workloads in Geo-Distributed Data Centers
von: Hogade, Ninad, et al.
Veröffentlicht: (2024)
von: Hogade, Ninad, et al.
Veröffentlicht: (2024)
Uncertainty-Aware Decarbonization for Datacenters
von: Li, Amy, et al.
Veröffentlicht: (2024)
von: Li, Amy, et al.
Veröffentlicht: (2024)
Characterization of Large Language Model Development in the Datacenter
von: Hu, Qinghao, et al.
Veröffentlicht: (2024)
von: Hu, Qinghao, et al.
Veröffentlicht: (2024)
Cost-aware Duration Prediction for Software Upgrades in Datacenters
von: Ding, Yi, et al.
Veröffentlicht: (2022)
von: Ding, Yi, et al.
Veröffentlicht: (2022)
Harnessing Your DRAM and SSD for Sustainable and Accessible LLM Inference with Mixed-Precision and Multi-level Caching
von: Peng, Jie, et al.
Veröffentlicht: (2024)
von: Peng, Jie, et al.
Veröffentlicht: (2024)
Enabling Efficient Serverless Inference Serving for LLM (Large Language Model) in the Cloud
von: Ghosh, Himel
Veröffentlicht: (2024)
von: Ghosh, Himel
Veröffentlicht: (2024)
Local-Cloud Inference Offloading for LLMs in Multi-Modal, Multi-Task, Multi-Dialogue Settings
von: Yuan, Liangqi, et al.
Veröffentlicht: (2025)
von: Yuan, Liangqi, et al.
Veröffentlicht: (2025)
Understanding and Improving Communication Performance in Multi-node LLM Inference
von: Singhania, Prajwal, et al.
Veröffentlicht: (2025)
von: Singhania, Prajwal, et al.
Veröffentlicht: (2025)
OpenG2G: A Simulation Platform for AI Datacenter-Grid Runtime Coordination
von: Chung, Jae-Won, et al.
Veröffentlicht: (2026)
von: Chung, Jae-Won, et al.
Veröffentlicht: (2026)
Making MoE-based LLM Inference Resilient with Tarragon
von: Zhang, Songyu, et al.
Veröffentlicht: (2026)
von: Zhang, Songyu, et al.
Veröffentlicht: (2026)
LLM Inference at the Edge: Mobile, NPU, and GPU Performance Efficiency Trade-offs Under Sustained Load
von: Tummalapalli, Pranay, et al.
Veröffentlicht: (2026)
von: Tummalapalli, Pranay, et al.
Veröffentlicht: (2026)
Collaborative Speculative Inference for Efficient LLM Inference Serving
von: Gao, Luyao, et al.
Veröffentlicht: (2025)
von: Gao, Luyao, et al.
Veröffentlicht: (2025)
Seer: Online Context Learning for Fast Synchronous LLM Reinforcement Learning
von: Qin, Ruoyu, et al.
Veröffentlicht: (2025)
von: Qin, Ruoyu, et al.
Veröffentlicht: (2025)
MAS-H2: A Hierarchical Multi-Agent System for Holistic Cloud-Native Autoscaling
von: Hamzeh, Hamed, et al.
Veröffentlicht: (2026)
von: Hamzeh, Hamed, et al.
Veröffentlicht: (2026)
AirFed: A Federated Graph-Enhanced Multi-Agent Reinforcement Learning Framework for Multi-UAV Cooperative Mobile Edge Computing
von: Wang, Zhiyu, et al.
Veröffentlicht: (2025)
von: Wang, Zhiyu, et al.
Veröffentlicht: (2025)
History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL
von: He, Jingkai, et al.
Veröffentlicht: (2025)
von: He, Jingkai, et al.
Veröffentlicht: (2025)
Improving the End-to-End Efficiency of Offline Inference for Multi-LLM Applications Based on Sampling and Simulation
von: Fang, Jingzhi, et al.
Veröffentlicht: (2025)
von: Fang, Jingzhi, et al.
Veröffentlicht: (2025)
Floe: Federated Specialization for Real-Time LLM-SLM Inference
von: Tian, Chunlin, et al.
Veröffentlicht: (2026)
von: Tian, Chunlin, et al.
Veröffentlicht: (2026)
Serving Compound Inference Systems on Datacenter GPUs
von: Devata, Sriram, et al.
Veröffentlicht: (2026)
von: Devata, Sriram, et al.
Veröffentlicht: (2026)
Adaptive, Efficient and Fair Resource Allocation in Cloud Datacenters leveraging Weighted A3C Deep Reinforcement Learning
von: Kumari, Suchi, et al.
Veröffentlicht: (2025)
von: Kumari, Suchi, et al.
Veröffentlicht: (2025)
Pie: Pooling CPU Memory for LLM Inference
von: Xu, Yi, et al.
Veröffentlicht: (2024)
von: Xu, Yi, et al.
Veröffentlicht: (2024)
STAR: Decode-Phase Rescheduling for LLM Inference
von: Wang, Zhibin, et al.
Veröffentlicht: (2025)
von: Wang, Zhibin, et al.
Veröffentlicht: (2025)
Optimal Scheduling Algorithms for LLM Inference: Theory and Practice
von: Bari, Agrim, et al.
Veröffentlicht: (2025)
von: Bari, Agrim, et al.
Veröffentlicht: (2025)
Autonomous Resource Management in Microservice Systems via Reinforcement Learning
von: Zou, Yujun, et al.
Veröffentlicht: (2025)
von: Zou, Yujun, et al.
Veröffentlicht: (2025)
Cost-Efficient LLM Serving in the Cloud: VM Selection with KV Cache Offloading
von: Kim, Kihyun, et al.
Veröffentlicht: (2025)
von: Kim, Kihyun, et al.
Veröffentlicht: (2025)
Learning to Shard: RL for Co-optimizing the Parallelism Degrees and Per-operator Sharding Dimensions in Distributed LLM Inference
von: Yin, Ruokai, et al.
Veröffentlicht: (2025)
von: Yin, Ruokai, et al.
Veröffentlicht: (2025)
ExeGPT: Constraint-Aware Resource Scheduling for LLM Inference
von: Oh, Hyungjun, et al.
Veröffentlicht: (2024)
von: Oh, Hyungjun, et al.
Veröffentlicht: (2024)
FDC: Fast KV Dimensionality Compression for Efficient LLM Inference
von: Zhang, Zeyu, et al.
Veröffentlicht: (2024)
von: Zhang, Zeyu, et al.
Veröffentlicht: (2024)
PecSched: Preemptive and Efficient Cluster Scheduling for LLM Inference
von: Zhang, Zeyu, et al.
Veröffentlicht: (2024)
von: Zhang, Zeyu, et al.
Veröffentlicht: (2024)
Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve
von: Agrawal, Amey, et al.
Veröffentlicht: (2024)
von: Agrawal, Amey, et al.
Veröffentlicht: (2024)
WarmServe: Enabling One-for-Many GPU Prewarming for Multi-LLM Serving
von: Lou, Chiheng, et al.
Veröffentlicht: (2025)
von: Lou, Chiheng, et al.
Veröffentlicht: (2025)
Challenging GPU Dominance: When CPUs Outperform for On-Device LLM Inference
von: Zhang, Haolin, et al.
Veröffentlicht: (2025)
von: Zhang, Haolin, et al.
Veröffentlicht: (2025)
ServerlessLLM: Low-Latency Serverless Inference for Large Language Models
von: Fu, Yao, et al.
Veröffentlicht: (2024)
von: Fu, Yao, et al.
Veröffentlicht: (2024)
TokenWeave: Efficient Compute-Communication Overlap for Distributed LLM Inference
von: Gond, Raja, et al.
Veröffentlicht: (2025)
von: Gond, Raja, et al.
Veröffentlicht: (2025)
Split CNN Inference on Networked Microcontrollers
von: Lu, Junyu, et al.
Veröffentlicht: (2026)
von: Lu, Junyu, et al.
Veröffentlicht: (2026)
CE-CoLLM: Efficient and Adaptive Large Language Models Through Cloud-Edge Collaboration
von: Jin, Hongpeng, et al.
Veröffentlicht: (2024)
von: Jin, Hongpeng, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Sustainable Carbon-Aware and Water-Efficient LLM Scheduling in Geo-Distributed Cloud Datacenters
von: Moore, Hayden, et al.
Veröffentlicht: (2025) -
Sustainable Graph Analytics Workload Scheduling with Evolutionary Reinforcement Learning in Edge-Cloud Systems
von: Ramicetty, P., et al.
Veröffentlicht: (2026) -
A Framework for SLO, Carbon, and Wastewater-Aware Sustainable FaaS Cloud Platform Management
von: Qi, Sirui, et al.
Veröffentlicht: (2024) -
CASA: A Framework for SLO and Carbon-Aware Autoscaling and Scheduling in Serverless Cloud Computing
von: Qi, S., et al.
Veröffentlicht: (2024) -
Game-Theoretic Deep Reinforcement Learning to Minimize Carbon Emissions and Energy Costs for AI Inference Workloads in Geo-Distributed Data Centers
von: Hogade, Ninad, et al.
Veröffentlicht: (2024)