Temporal-Aware GPU Resource Allocation for Distributed LLM Inference via Reinforcement Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Du, Chengze, Yu, Zhiwei, Xu, Heng, Wang, Haojie, liu, Bo, Li, Jialong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RailS: Load Balancing for All-to-All Communication in Distributed Mixture-of-Experts Training
by: Xu, Heng, et al.
Published: (2025)
by: Xu, Heng, et al.
Published: (2025)
Optimizing Resource Allocation for Geographically-Distributed Inference by Large Language Models
by: Sun, Tingyang, et al.
Published: (2025)
by: Sun, Tingyang, et al.
Published: (2025)
RouterWise: Joint Resource Allocation and Routing for Latency-Aware Multi-Model LLM Serving
by: Kasnavieh, Hossein Hosseini, et al.
Published: (2026)
by: Kasnavieh, Hossein Hosseini, et al.
Published: (2026)
An Auction-Based Mechanism for Optimal Task Allocation and Resource Aware Containerization
by: kumar, Ramakant
Published: (2026)
by: kumar, Ramakant
Published: (2026)
Resource Allocation Driven by Large Models in Future Semantic-Aware Networks
by: Zhang, Haijun, et al.
Published: (2025)
by: Zhang, Haijun, et al.
Published: (2025)
HALO: Semantic-Aware Distributed LLM Inference in Lossy Edge Network
by: Zheng, Peirong, et al.
Published: (2026)
by: Zheng, Peirong, et al.
Published: (2026)
Green-LLM: Optimal Workload Allocation for Environmentally-Aware Distributed Inference
by: Cheng, Jiaming, et al.
Published: (2025)
by: Cheng, Jiaming, et al.
Published: (2025)
Semantic-Aware LLM Orchestration for Proactive Resource Management in Predictive Digital Twin Vehicular Networks
by: Ahmadpanah, Seyed Hossein
Published: (2025)
by: Ahmadpanah, Seyed Hossein
Published: (2025)
An Online Fragmentation-Aware GPU Scheduler for Multi-Tenant MIG-based Clouds
by: Zambianco, Marco, et al.
Published: (2025)
by: Zambianco, Marco, et al.
Published: (2025)
PerLLM: Personalized Inference Scheduling with Edge-Cloud Collaboration for Diverse LLM Services
by: Yang, Zheming, et al.
Published: (2024)
by: Yang, Zheming, et al.
Published: (2024)
FAST: An Efficient Scheduler for All-to-All GPU Communication
by: Lei, Yiran, et al.
Published: (2025)
by: Lei, Yiran, et al.
Published: (2025)
Carbon-Aware Temporal Data Transfer Scheduling Across Cloud Datacenters
by: Rodrigues, Elvis, et al.
Published: (2025)
by: Rodrigues, Elvis, et al.
Published: (2025)
MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services
by: Tang, Lingfeng, et al.
Published: (2025)
by: Tang, Lingfeng, et al.
Published: (2025)
Trust-Aware Routing for Distributed Generative AI Inference at the Edge
by: Nguyen, Chanh, et al.
Published: (2026)
by: Nguyen, Chanh, et al.
Published: (2026)
A Task Decomposition and Planning Framework for Efficient LLM Inference in AI-Enabled WiFi-Offload Networks
by: Han, Mingqi, et al.
Published: (2026)
by: Han, Mingqi, et al.
Published: (2026)
Dynamic Hierarchical Birkhoff-von Neumann Decomposition for All-to-All GPU Communication
by: Wu, Yen-Chieh, et al.
Published: (2026)
by: Wu, Yen-Chieh, et al.
Published: (2026)
Recursive Offloading for LLM Serving in Multi-tier Networks
by: Wu, Zhiyuan, et al.
Published: (2025)
by: Wu, Zhiyuan, et al.
Published: (2025)
From Skew to Symmetry: Node-Interconnect Multi-Path Balancing with Execution-time Planning for Modern GPU Clusters
by: Yao, Jinghan, et al.
Published: (2026)
by: Yao, Jinghan, et al.
Published: (2026)
QoS-Aware Load Balancing in the Computing Continuum via Multi-Player Bandits
by: Čilić, Ivan, et al.
Published: (2025)
by: Čilić, Ivan, et al.
Published: (2025)
Varuna: Enabling Failure-Type Aware RDMA Failover
by: Wang, Xiaoyang, et al.
Published: (2026)
by: Wang, Xiaoyang, et al.
Published: (2026)
D-LoRa: a Distributed Parameter Adaptation Scheme for LoRa Network
by: Wang, Ruiqi, et al.
Published: (2025)
by: Wang, Ruiqi, et al.
Published: (2025)
Deep-Reinforcement-Learning-Based AoI-Aware Resource Allocation for RIS-Aided IoV Networks
by: Qi, Kangwei, et al.
Published: (2024)
by: Qi, Kangwei, et al.
Published: (2024)
Surviving the Edge: Federated Learning under Networking and Resource Constraints
by: Mwanje, Mike, et al.
Published: (2026)
by: Mwanje, Mike, et al.
Published: (2026)
A Survey on Resource Management in Joint Communication and Computing-Embedded SAGIN
by: Chen, Qian, et al.
Published: (2024)
by: Chen, Qian, et al.
Published: (2024)
Jupiter: Fast and Resource-Efficient Collaborative Inference of Generative LLMs on Edge Devices
by: Ye, Shengyuan, et al.
Published: (2025)
by: Ye, Shengyuan, et al.
Published: (2025)
Causal Inference for Quantifying Noisy Neighbor Effects in Multi-Tenant Cloud Environments
by: Schiavo, Philipe S., et al.
Published: (2026)
by: Schiavo, Philipe S., et al.
Published: (2026)
AllReduce Scheduling with Hierarchical Deep Reinforcement Learning
by: Wei, Yufan, et al.
Published: (2025)
by: Wei, Yufan, et al.
Published: (2025)
Toward Edge General Intelligence with Multiple-Large Language Model (Multi-LLM): Architecture, Trust, and Orchestration
by: Luo, Haoxiang, et al.
Published: (2025)
by: Luo, Haoxiang, et al.
Published: (2025)
Multi-stage Flow Scheduling for LLM Serving
by: Sun, Yijun, et al.
Published: (2026)
by: Sun, Yijun, et al.
Published: (2026)
Legible Consensus: Topology-Aware Quorum Geometry for Asymmetric Networks
by: Mason, Tony
Published: (2026)
by: Mason, Tony
Published: (2026)
Contention-Aware Microservice Deployment in Collaborative Mobile Edge Networks
by: Ge, Xinlei, et al.
Published: (2024)
by: Ge, Xinlei, et al.
Published: (2024)
SARS: A Resource Selection Algorithm for Autonomous Driving Tasks in Heterogeneous Mobile Edge Computing
by: Zakerian, Reza, et al.
Published: (2024)
by: Zakerian, Reza, et al.
Published: (2024)
Network Anomaly Detection in Distributed Edge Computing Infrastructure
by: Marfo, William, et al.
Published: (2025)
by: Marfo, William, et al.
Published: (2025)
Future Resource Bank for ISAC: Achieving Fast and Stable Win-Win Matching for Both Individuals and Coalitions
by: Qi, Houyi, et al.
Published: (2025)
by: Qi, Houyi, et al.
Published: (2025)
A Uniqueness Theorem for Distributed Computation under Physical Constraint
by: Ren, Zhiyuan, et al.
Published: (2025)
by: Ren, Zhiyuan, et al.
Published: (2025)
Policy Design in Zero-Trust Distributed Networks: Challenges and Solutions
by: Sandjaja, Fannya R., et al.
Published: (2025)
by: Sandjaja, Fannya R., et al.
Published: (2025)
Risk-Aware and Stable Edge Server Selection Under Network Latency SLOs
by: Liyanage, Mohan, et al.
Published: (2026)
by: Liyanage, Mohan, et al.
Published: (2026)
Accelerating Stable Matching between Workers and Spatial-Temporal Tasks for Dynamic MCS: A Stagewise Service Trading Approach
by: Qi, Houyi, et al.
Published: (2025)
by: Qi, Houyi, et al.
Published: (2025)
Q-adaptive: A Multi-Agent Reinforcement Learning Based Routing on Dragonfly Network
by: Kang, Yao, et al.
Published: (2024)
by: Kang, Yao, et al.
Published: (2024)
DistriFS: A Platform and User Agnostic Approach to File Distribution
by: Boesch, Julian
Published: (2024)
by: Boesch, Julian
Published: (2024)
Similar Items
-
RailS: Load Balancing for All-to-All Communication in Distributed Mixture-of-Experts Training
by: Xu, Heng, et al.
Published: (2025) -
Optimizing Resource Allocation for Geographically-Distributed Inference by Large Language Models
by: Sun, Tingyang, et al.
Published: (2025) -
RouterWise: Joint Resource Allocation and Routing for Latency-Aware Multi-Model LLM Serving
by: Kasnavieh, Hossein Hosseini, et al.
Published: (2026) -
An Auction-Based Mechanism for Optimal Task Allocation and Resource Aware Containerization
by: kumar, Ramakant
Published: (2026) -
Resource Allocation Driven by Large Models in Future Semantic-Aware Networks
by: Zhang, Haijun, et al.
Published: (2025)