RouterWise: Joint Resource Allocation and Routing for Latency-Aware Multi-Model LLM Serving
Fuente:
arXiv
Saved in:
| Main Authors: | Kasnavieh, Hossein Hosseini, Leckie, Christopher, Toosi, Adel N. |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Semantic-Aware LLM Orchestration for Proactive Resource Management in Predictive Digital Twin Vehicular Networks
by: Ahmadpanah, Seyed Hossein
Published: (2025)
by: Ahmadpanah, Seyed Hossein
Published: (2025)
Temporal-Aware GPU Resource Allocation for Distributed LLM Inference via Reinforcement Learning
by: Du, Chengze, et al.
Published: (2025)
by: Du, Chengze, et al.
Published: (2025)
Multi-stage Flow Scheduling for LLM Serving
by: Sun, Yijun, et al.
Published: (2026)
by: Sun, Yijun, et al.
Published: (2026)
Recursive Offloading for LLM Serving in Multi-tier Networks
by: Wu, Zhiyuan, et al.
Published: (2025)
by: Wu, Zhiyuan, et al.
Published: (2025)
An Auction-Based Mechanism for Optimal Task Allocation and Resource Aware Containerization
by: kumar, Ramakant
Published: (2026)
by: kumar, Ramakant
Published: (2026)
Resource Allocation Driven by Large Models in Future Semantic-Aware Networks
by: Zhang, Haijun, et al.
Published: (2025)
by: Zhang, Haijun, et al.
Published: (2025)
Risk-Aware and Stable Edge Server Selection Under Network Latency SLOs
by: Liyanage, Mohan, et al.
Published: (2026)
by: Liyanage, Mohan, et al.
Published: (2026)
XWind: A Cross-site Router for Large Language Model Inference Serving at Renewable Energy Farms
by: Reddy, Tella Rajashekhar, et al.
Published: (2026)
by: Reddy, Tella Rajashekhar, et al.
Published: (2026)
GORGO: Maximizing KV-Cache Reuse While Minimizing Network Latency in Cross-Region LLM Load Balancing
by: Toniolo, Alessio Ricci, et al.
Published: (2026)
by: Toniolo, Alessio Ricci, et al.
Published: (2026)
A Survey on Resource Management in Joint Communication and Computing-Embedded SAGIN
by: Chen, Qian, et al.
Published: (2024)
by: Chen, Qian, et al.
Published: (2024)
CRAFT: Latency and Cost-Aware Genetic-Based Framework for Node Placement in Edge-Fog Environments
by: Mahdizadeh, Soheil, et al.
Published: (2025)
by: Mahdizadeh, Soheil, et al.
Published: (2025)
Short-circuiting Rings for Low-Latency AllReduce
by: Hammer, Sarah-Michelle, et al.
Published: (2025)
by: Hammer, Sarah-Michelle, et al.
Published: (2025)
Solving AI Foundational Model Latency with Telco Infrastructure
by: Barros, Sebastian
Published: (2025)
by: Barros, Sebastian
Published: (2025)
KVServe: Service-Aware KV Cache Compression for Communication-Efficient Disaggregated LLM Serving
by: Liu, Zedong, et al.
Published: (2026)
by: Liu, Zedong, et al.
Published: (2026)
Trivance: Latency-Optimal AllReduce by Shortcutting Multiport Networks
by: Juerss, Anton, et al.
Published: (2026)
by: Juerss, Anton, et al.
Published: (2026)
Q-adaptive: A Multi-Agent Reinforcement Learning Based Routing on Dragonfly Network
by: Kang, Yao, et al.
Published: (2024)
by: Kang, Yao, et al.
Published: (2024)
Joint Network Slicing, Routing, and In-Network Computing for Energy-Efficient 6G
by: Sasan, Zeinab, et al.
Published: (2024)
by: Sasan, Zeinab, et al.
Published: (2024)
Lightweight Latency Prediction Scheme for Edge Applications: A Rational Modelling Approach
by: Liyanage, Mohan, et al.
Published: (2025)
by: Liyanage, Mohan, et al.
Published: (2025)
Low-Latency Video Conferencing via Optimized Packet Routing and Reordering
by: Xiao, Yao, et al.
Published: (2023)
by: Xiao, Yao, et al.
Published: (2023)
COREC: Concurrent Non-Blocking Single-Queue Receive Driver for Low Latency Networking
by: Faltelli, Marco, et al.
Published: (2024)
by: Faltelli, Marco, et al.
Published: (2024)
Study of Workload Interference with Intelligent Routing on Dragonfly
by: Kang, Yao, et al.
Published: (2024)
by: Kang, Yao, et al.
Published: (2024)
Distributed Simulation for Digital Twins of Large-Scale Real-World DiffServ-Based Networks
by: Huang, Zhuoyao, et al.
Published: (2024)
by: Huang, Zhuoyao, et al.
Published: (2024)
1.5 Million Messages Per Second on 3 Machines: Benchmarking and Latency Optimization of Apache Pulsar at Enterprise Scale
by: Mukkolakkal, Muhamed Ramees Cheriya
Published: (2026)
by: Mukkolakkal, Muhamed Ramees Cheriya
Published: (2026)
QoS-Aware Load Balancing in the Computing Continuum via Multi-Player Bandits
by: Čilić, Ivan, et al.
Published: (2025)
by: Čilić, Ivan, et al.
Published: (2025)
An Online Fragmentation-Aware GPU Scheduler for Multi-Tenant MIG-based Clouds
by: Zambianco, Marco, et al.
Published: (2025)
by: Zambianco, Marco, et al.
Published: (2025)
An Open-Source Fast Parallel Routing Approach for Commercial FPGAs
by: Zang, Xinshi, et al.
Published: (2024)
by: Zang, Xinshi, et al.
Published: (2024)
Reliable and Resilient Collective Communication Library for LLM Training and Serving
by: Wang, Wei, et al.
Published: (2025)
by: Wang, Wei, et al.
Published: (2025)
Leveraging InfiniBand Controller to Configure Deadlock-Free Routing Engines for Dragonflies
by: Maglione-Mathey, German, et al.
Published: (2025)
by: Maglione-Mathey, German, et al.
Published: (2025)
Toward Edge General Intelligence with Multiple-Large Language Model (Multi-LLM): Architecture, Trust, and Orchestration
by: Luo, Haoxiang, et al.
Published: (2025)
by: Luo, Haoxiang, et al.
Published: (2025)
Optimizing Resource Allocation for Geographically-Distributed Inference by Large Language Models
by: Sun, Tingyang, et al.
Published: (2025)
by: Sun, Tingyang, et al.
Published: (2025)
A Lightweight and Scalable Design of Segment Routing in Broadband LEO Constellations Using Landmark-Based Skeleton Graphs
by: Hu, Menglan, et al.
Published: (2024)
by: Hu, Menglan, et al.
Published: (2024)
FogROS2-PLR: Probabilistic Latency-Reliability For Cloud Robotics
by: Chen, Kaiyuan, et al.
Published: (2024)
by: Chen, Kaiyuan, et al.
Published: (2024)
Joint wireless and computing resource management with optimal slice selection in in-network-edge metaverse system
by: Rashid, Sulaiman Muhammad, et al.
Published: (2024)
by: Rashid, Sulaiman Muhammad, et al.
Published: (2024)
LIMO: Load-balanced Offloading with MAPE and Particle Swarm Optimization in Mobile Fog Networks
by: Seraj, Yasaman, et al.
Published: (2024)
by: Seraj, Yasaman, et al.
Published: (2024)
Surviving the Edge: Federated Learning under Networking and Resource Constraints
by: Mwanje, Mike, et al.
Published: (2026)
by: Mwanje, Mike, et al.
Published: (2026)
Trust-Aware Routing for Distributed Generative AI Inference at the Edge
by: Nguyen, Chanh, et al.
Published: (2026)
by: Nguyen, Chanh, et al.
Published: (2026)
PerLLM: Personalized Inference Scheduling with Edge-Cloud Collaboration for Diverse LLM Services
by: Yang, Zheming, et al.
Published: (2024)
by: Yang, Zheming, et al.
Published: (2024)
Varuna: Enabling Failure-Type Aware RDMA Failover
by: Wang, Xiaoyang, et al.
Published: (2026)
by: Wang, Xiaoyang, et al.
Published: (2026)
Legible Consensus: Topology-Aware Quorum Geometry for Asymmetric Networks
by: Mason, Tony
Published: (2026)
by: Mason, Tony
Published: (2026)
Contention-Aware Microservice Deployment in Collaborative Mobile Edge Networks
by: Ge, Xinlei, et al.
Published: (2024)
by: Ge, Xinlei, et al.
Published: (2024)
Similar Items
-
Semantic-Aware LLM Orchestration for Proactive Resource Management in Predictive Digital Twin Vehicular Networks
by: Ahmadpanah, Seyed Hossein
Published: (2025) -
Temporal-Aware GPU Resource Allocation for Distributed LLM Inference via Reinforcement Learning
by: Du, Chengze, et al.
Published: (2025) -
Multi-stage Flow Scheduling for LLM Serving
by: Sun, Yijun, et al.
Published: (2026) -
Recursive Offloading for LLM Serving in Multi-tier Networks
by: Wu, Zhiyuan, et al.
Published: (2025) -
An Auction-Based Mechanism for Optimal Task Allocation and Resource Aware Containerization
by: kumar, Ramakant
Published: (2026)