GORGO: Maximizing KV-Cache Reuse While Minimizing Network Latency in Cross-Region LLM Load Balancing
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Toniolo, Alessio Ricci, Dinesh, Abinaya, Thorstenson, Rome |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Deduplicator: When Computation Reuse Meets Load Balancing at the Network Edge
von: Azad, Md Washik Al, et al.
Veröffentlicht: (2024)
von: Azad, Md Washik Al, et al.
Veröffentlicht: (2024)
Local Rendezvous Hashing: Bounded Loads and Minimal Churn via Cache-Local Candidates
von: Guan, Yongjie
Veröffentlicht: (2025)
von: Guan, Yongjie
Veröffentlicht: (2025)
Adaptive KV Cache Reuse for Fast Long-Context LLM Serving
von: li, Fei, et al.
Veröffentlicht: (2026)
von: li, Fei, et al.
Veröffentlicht: (2026)
KVServe: Service-Aware KV Cache Compression for Communication-Efficient Disaggregated LLM Serving
von: Liu, Zedong, et al.
Veröffentlicht: (2026)
von: Liu, Zedong, et al.
Veröffentlicht: (2026)
Optimal Oblivious Load-Balancing for Sparse Traffic in Large-Scale Satellite Networks
von: Ramakanth, Rudrapatna Vallabh, et al.
Veröffentlicht: (2026)
von: Ramakanth, Rudrapatna Vallabh, et al.
Veröffentlicht: (2026)
RailS: Load Balancing for All-to-All Communication in Distributed Mixture-of-Experts Training
von: Xu, Heng, et al.
Veröffentlicht: (2025)
von: Xu, Heng, et al.
Veröffentlicht: (2025)
QoS-Aware Load Balancing in the Computing Continuum via Multi-Player Bandits
von: Čilić, Ivan, et al.
Veröffentlicht: (2025)
von: Čilić, Ivan, et al.
Veröffentlicht: (2025)
Ethereal: Divide and Conquer Network Load Balancing in Large-Scale Distributed Training
von: Addanki, Vamsi, et al.
Veröffentlicht: (2024)
von: Addanki, Vamsi, et al.
Veröffentlicht: (2024)
RouterWise: Joint Resource Allocation and Routing for Latency-Aware Multi-Model LLM Serving
von: Kasnavieh, Hossein Hosseini, et al.
Veröffentlicht: (2026)
von: Kasnavieh, Hossein Hosseini, et al.
Veröffentlicht: (2026)
Short-circuiting Rings for Low-Latency AllReduce
von: Hammer, Sarah-Michelle, et al.
Veröffentlicht: (2025)
von: Hammer, Sarah-Michelle, et al.
Veröffentlicht: (2025)
Solving AI Foundational Model Latency with Telco Infrastructure
von: Barros, Sebastian
Veröffentlicht: (2025)
von: Barros, Sebastian
Veröffentlicht: (2025)
Trivance: Latency-Optimal AllReduce by Shortcutting Multiport Networks
von: Juerss, Anton, et al.
Veröffentlicht: (2026)
von: Juerss, Anton, et al.
Veröffentlicht: (2026)
Risk-Aware and Stable Edge Server Selection Under Network Latency SLOs
von: Liyanage, Mohan, et al.
Veröffentlicht: (2026)
von: Liyanage, Mohan, et al.
Veröffentlicht: (2026)
Lightweight Latency Prediction Scheme for Edge Applications: A Rational Modelling Approach
von: Liyanage, Mohan, et al.
Veröffentlicht: (2025)
von: Liyanage, Mohan, et al.
Veröffentlicht: (2025)
COREC: Concurrent Non-Blocking Single-Queue Receive Driver for Low Latency Networking
von: Faltelli, Marco, et al.
Veröffentlicht: (2024)
von: Faltelli, Marco, et al.
Veröffentlicht: (2024)
LOAM: Low-latency Communication, Caching, and Computation Placement in Data-Intensive Computing Networks
von: Zhang, Jinkun, et al.
Veröffentlicht: (2024)
von: Zhang, Jinkun, et al.
Veröffentlicht: (2024)
A Hybrid Approach to Monitor Context Parameters for Optimising Caching for Context-Aware IoT Applications
von: Manchanda, Ashish, et al.
Veröffentlicht: (2024)
von: Manchanda, Ashish, et al.
Veröffentlicht: (2024)
1.5 Million Messages Per Second on 3 Machines: Benchmarking and Latency Optimization of Apache Pulsar at Enterprise Scale
von: Mukkolakkal, Muhamed Ramees Cheriya
Veröffentlicht: (2026)
von: Mukkolakkal, Muhamed Ramees Cheriya
Veröffentlicht: (2026)
CRAFT: Latency and Cost-Aware Genetic-Based Framework for Node Placement in Edge-Fog Environments
von: Mahdizadeh, Soheil, et al.
Veröffentlicht: (2025)
von: Mahdizadeh, Soheil, et al.
Veröffentlicht: (2025)
LIMO: Load-balanced Offloading with MAPE and Particle Swarm Optimization in Mobile Fog Networks
von: Seraj, Yasaman, et al.
Veröffentlicht: (2024)
von: Seraj, Yasaman, et al.
Veröffentlicht: (2024)
Heterogeneity-aware P2P Wireless Energy Transfer for Balanced Energy Distribution
von: Ojha, Tamoghna, et al.
Veröffentlicht: (2022)
von: Ojha, Tamoghna, et al.
Veröffentlicht: (2022)
FogROS2-PLR: Probabilistic Latency-Reliability For Cloud Robotics
von: Chen, Kaiyuan, et al.
Veröffentlicht: (2024)
von: Chen, Kaiyuan, et al.
Veröffentlicht: (2024)
GENIO: Synergizing Edge Computing with Optical Network Infrastructures
von: Cesarano, Carmine, et al.
Veröffentlicht: (2025)
von: Cesarano, Carmine, et al.
Veröffentlicht: (2025)
PerLLM: Personalized Inference Scheduling with Edge-Cloud Collaboration for Diverse LLM Services
von: Yang, Zheming, et al.
Veröffentlicht: (2024)
von: Yang, Zheming, et al.
Veröffentlicht: (2024)
From Skew to Symmetry: Node-Interconnect Multi-Path Balancing with Execution-time Planning for Modern GPU Clusters
von: Yao, Jinghan, et al.
Veröffentlicht: (2026)
von: Yao, Jinghan, et al.
Veröffentlicht: (2026)
SlimCaching: Edge Caching of Mixture-of-Experts for Distributed Inference
von: Chen, Qian, et al.
Veröffentlicht: (2025)
von: Chen, Qian, et al.
Veröffentlicht: (2025)
Move the Query, Not the Cache: Characterizing Cross-Instance Latent Attention Redistribution Across GPU Fabrics
von: Ma, Bole, et al.
Veröffentlicht: (2026)
von: Ma, Bole, et al.
Veröffentlicht: (2026)
Multi-stage Flow Scheduling for LLM Serving
von: Sun, Yijun, et al.
Veröffentlicht: (2026)
von: Sun, Yijun, et al.
Veröffentlicht: (2026)
Recursive Offloading for LLM Serving in Multi-tier Networks
von: Wu, Zhiyuan, et al.
Veröffentlicht: (2025)
von: Wu, Zhiyuan, et al.
Veröffentlicht: (2025)
Decentralized Stratified Sampling for Low-Latency Approximate Geospatial Data Stream Processing in Edge-Cloud Architectures
von: Jawarneh, Isam Mashhour Al, et al.
Veröffentlicht: (2026)
von: Jawarneh, Isam Mashhour Al, et al.
Veröffentlicht: (2026)
Adaptive Multi-Objective Tiered Storage Configuration for KV Cache in LLM Service
von: Zheng, Xianzhe, et al.
Veröffentlicht: (2026)
von: Zheng, Xianzhe, et al.
Veröffentlicht: (2026)
CCL-Bench 1.0: A Trace-Based Benchmark for LLM Infrastructure
von: Ding, Eric, et al.
Veröffentlicht: (2026)
von: Ding, Eric, et al.
Veröffentlicht: (2026)
Temporal-Aware GPU Resource Allocation for Distributed LLM Inference via Reinforcement Learning
von: Du, Chengze, et al.
Veröffentlicht: (2025)
von: Du, Chengze, et al.
Veröffentlicht: (2025)
BlockSDN: Towards a High-Performance Blockchain via Software-Defined Cross Networking optimization
von: Jia, Wenyang, et al.
Veröffentlicht: (2025)
von: Jia, Wenyang, et al.
Veröffentlicht: (2025)
Semantic-Aware LLM Orchestration for Proactive Resource Management in Predictive Digital Twin Vehicular Networks
von: Ahmadpanah, Seyed Hossein
Veröffentlicht: (2025)
von: Ahmadpanah, Seyed Hossein
Veröffentlicht: (2025)
Toward Edge General Intelligence with Multiple-Large Language Model (Multi-LLM): Architecture, Trust, and Orchestration
von: Luo, Haoxiang, et al.
Veröffentlicht: (2025)
von: Luo, Haoxiang, et al.
Veröffentlicht: (2025)
A Task Decomposition and Planning Framework for Efficient LLM Inference in AI-Enabled WiFi-Offload Networks
von: Han, Mingqi, et al.
Veröffentlicht: (2026)
von: Han, Mingqi, et al.
Veröffentlicht: (2026)
SAKURAONE: An Open Ethernet-Based AI HPC System and Its Observed Workload Dynamics in a Single-Tenant LLM Development Environment
von: Konishi, Fumikazu, et al.
Veröffentlicht: (2026)
von: Konishi, Fumikazu, et al.
Veröffentlicht: (2026)
Optimizing Split Learning Latency in TinyML-Based IoT Systems
von: Jenhani, Zied, et al.
Veröffentlicht: (2025)
von: Jenhani, Zied, et al.
Veröffentlicht: (2025)
Revisiting Cache Freshness for Emerging Real-Time Applications
von: Mao, Ziming, et al.
Veröffentlicht: (2024)
von: Mao, Ziming, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Deduplicator: When Computation Reuse Meets Load Balancing at the Network Edge
von: Azad, Md Washik Al, et al.
Veröffentlicht: (2024) -
Local Rendezvous Hashing: Bounded Loads and Minimal Churn via Cache-Local Candidates
von: Guan, Yongjie
Veröffentlicht: (2025) -
Adaptive KV Cache Reuse for Fast Long-Context LLM Serving
von: li, Fei, et al.
Veröffentlicht: (2026) -
KVServe: Service-Aware KV Cache Compression for Communication-Efficient Disaggregated LLM Serving
von: Liu, Zedong, et al.
Veröffentlicht: (2026) -
Optimal Oblivious Load-Balancing for Sparse Traffic in Large-Scale Satellite Networks
von: Ramakanth, Rudrapatna Vallabh, et al.
Veröffentlicht: (2026)