BeLLMan: Controlling LLM Congestion
Fuente:
arXiv
Salvato in:
| Autori principali: | Reddy, Tella Rajashekhar, Deshmukh, Atharva, Tandon, Karan, Gandhi, Rohan, Parayil, Anjaly, Bhattacherjee, Debopam |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
XWind: A Cross-site Router for Large Language Model Inference Serving at Renewable Energy Farms
di: Reddy, Tella Rajashekhar, et al.
Pubblicazione: (2026)
di: Reddy, Tella Rajashekhar, et al.
Pubblicazione: (2026)
AI Greenferencing: Routing AI Inferencing to Green Modular Data Centers with Heron
di: Reddy, Tella Rajashekhar, et al.
Pubblicazione: (2025)
di: Reddy, Tella Rajashekhar, et al.
Pubblicazione: (2025)
Improving training time and GPU utilization in geo-distributed language model training
di: Palak, et al.
Pubblicazione: (2024)
di: Palak, et al.
Pubblicazione: (2024)
Local Fast Rerouting with Low Congestion: A Randomized Approach
di: Bankhamer, Gregor, et al.
Pubblicazione: (2020)
di: Bankhamer, Gregor, et al.
Pubblicazione: (2020)
MLTCP: Congestion Control for DNN Training
di: Rajasekaran, Sudarsanan, et al.
Pubblicazione: (2024)
di: Rajasekaran, Sudarsanan, et al.
Pubblicazione: (2024)
PerLLM: Personalized Inference Scheduling with Edge-Cloud Collaboration for Diverse LLM Services
di: Yang, Zheming, et al.
Pubblicazione: (2024)
di: Yang, Zheming, et al.
Pubblicazione: (2024)
Hypergraph based Multi-Party Payment Channel
di: Nainwal, Ayush, et al.
Pubblicazione: (2025)
di: Nainwal, Ayush, et al.
Pubblicazione: (2025)
Multi-stage Flow Scheduling for LLM Serving
di: Sun, Yijun, et al.
Pubblicazione: (2026)
di: Sun, Yijun, et al.
Pubblicazione: (2026)
Recursive Offloading for LLM Serving in Multi-tier Networks
di: Wu, Zhiyuan, et al.
Pubblicazione: (2025)
di: Wu, Zhiyuan, et al.
Pubblicazione: (2025)
SDNator is Not Another SDN Controller: Enabling Extensible Data-Driven Control in Cyber-Physical Systems
di: Lin, Y., et al.
Pubblicazione: (2026)
di: Lin, Y., et al.
Pubblicazione: (2026)
CCL-Bench 1.0: A Trace-Based Benchmark for LLM Infrastructure
di: Ding, Eric, et al.
Pubblicazione: (2026)
di: Ding, Eric, et al.
Pubblicazione: (2026)
Temporal-Aware GPU Resource Allocation for Distributed LLM Inference via Reinforcement Learning
di: Du, Chengze, et al.
Pubblicazione: (2025)
di: Du, Chengze, et al.
Pubblicazione: (2025)
Semantic-Aware LLM Orchestration for Proactive Resource Management in Predictive Digital Twin Vehicular Networks
di: Ahmadpanah, Seyed Hossein
Pubblicazione: (2025)
di: Ahmadpanah, Seyed Hossein
Pubblicazione: (2025)
RouterWise: Joint Resource Allocation and Routing for Latency-Aware Multi-Model LLM Serving
di: Kasnavieh, Hossein Hosseini, et al.
Pubblicazione: (2026)
di: Kasnavieh, Hossein Hosseini, et al.
Pubblicazione: (2026)
Toward Edge General Intelligence with Multiple-Large Language Model (Multi-LLM): Architecture, Trust, and Orchestration
di: Luo, Haoxiang, et al.
Pubblicazione: (2025)
di: Luo, Haoxiang, et al.
Pubblicazione: (2025)
A Task Decomposition and Planning Framework for Efficient LLM Inference in AI-Enabled WiFi-Offload Networks
di: Han, Mingqi, et al.
Pubblicazione: (2026)
di: Han, Mingqi, et al.
Pubblicazione: (2026)
GORGO: Maximizing KV-Cache Reuse While Minimizing Network Latency in Cross-Region LLM Load Balancing
di: Toniolo, Alessio Ricci, et al.
Pubblicazione: (2026)
di: Toniolo, Alessio Ricci, et al.
Pubblicazione: (2026)
Leveraging InfiniBand Controller to Configure Deadlock-Free Routing Engines for Dragonflies
di: Maglione-Mathey, German, et al.
Pubblicazione: (2025)
di: Maglione-Mathey, German, et al.
Pubblicazione: (2025)
SAKURAONE: An Open Ethernet-Based AI HPC System and Its Observed Workload Dynamics in a Single-Tenant LLM Development Environment
di: Konishi, Fumikazu, et al.
Pubblicazione: (2026)
di: Konishi, Fumikazu, et al.
Pubblicazione: (2026)
Omnichain Web: The Universal Framework for Streamlined Chain Abstraction and Cross-Layer Interaction
di: Gajera, Hardik, et al.
Pubblicazione: (2024)
di: Gajera, Hardik, et al.
Pubblicazione: (2024)
RailX: A Flexible, Scalable, and Low-Cost Network Architecture for Hyper-Scale LLM Training Systems
di: Feng, Yinxiao, et al.
Pubblicazione: (2025)
di: Feng, Yinxiao, et al.
Pubblicazione: (2025)
MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services
di: Tang, Lingfeng, et al.
Pubblicazione: (2025)
di: Tang, Lingfeng, et al.
Pubblicazione: (2025)
Trivance: Latency-Optimal AllReduce by Shortcutting Multiport Networks
di: Juerss, Anton, et al.
Pubblicazione: (2026)
di: Juerss, Anton, et al.
Pubblicazione: (2026)
A Uniqueness Theorem for Distributed Computation under Physical Constraint
di: Ren, Zhiyuan, et al.
Pubblicazione: (2025)
di: Ren, Zhiyuan, et al.
Pubblicazione: (2025)
Fast Multichannel Topology Discovery in Cognitive Radio Networks
di: Wang, Yung-Li, et al.
Pubblicazione: (2025)
di: Wang, Yung-Li, et al.
Pubblicazione: (2025)
Enabling Scalability in Asynchronous and Bidirectional Communication in LPWAN
di: Rahman, Mahbubur
Pubblicazione: (2025)
di: Rahman, Mahbubur
Pubblicazione: (2025)
Real-Time In-Network Machine Learning on P4-Programmable FPGA SmartNICs with Fixed-Point Arithmetic and Taylor
di: Sada, Mohammad Firas, et al.
Pubblicazione: (2025)
di: Sada, Mohammad Firas, et al.
Pubblicazione: (2025)
Performance Evaluation of Brokerless Messaging Libraries
di: La Corte, Lorenzo, et al.
Pubblicazione: (2025)
di: La Corte, Lorenzo, et al.
Pubblicazione: (2025)
Accelerating Stable Matching between Workers and Spatial-Temporal Tasks for Dynamic MCS: A Stagewise Service Trading Approach
di: Qi, Houyi, et al.
Pubblicazione: (2025)
di: Qi, Houyi, et al.
Pubblicazione: (2025)
An Auction-Based Mechanism for Optimal Task Allocation and Resource Aware Containerization
di: kumar, Ramakant
Pubblicazione: (2026)
di: kumar, Ramakant
Pubblicazione: (2026)
SCRamble: Adaptive Decentralized Overlay Construction for Blockchain Networks
di: Kolyvas, Evangelos, et al.
Pubblicazione: (2026)
di: Kolyvas, Evangelos, et al.
Pubblicazione: (2026)
Optimal Server Selection for Straggler Mitigation
di: Badita, Ajay, et al.
Pubblicazione: (2019)
di: Badita, Ajay, et al.
Pubblicazione: (2019)
A Density-Delay Law for Stable Event-Driven State Progression in Open Distributed Systems
di: Chen, Bin, et al.
Pubblicazione: (2026)
di: Chen, Bin, et al.
Pubblicazione: (2026)
Optimal Oblivious Load-Balancing for Sparse Traffic in Large-Scale Satellite Networks
di: Ramakanth, Rudrapatna Vallabh, et al.
Pubblicazione: (2026)
di: Ramakanth, Rudrapatna Vallabh, et al.
Pubblicazione: (2026)
ncsim: A Lightweight Simulator for Networked Edge Computing with Wireless Interference Modeling
di: Krishnamachari, Bhaskar, et al.
Pubblicazione: (2026)
di: Krishnamachari, Bhaskar, et al.
Pubblicazione: (2026)
COREC: Concurrent Non-Blocking Single-Queue Receive Driver for Low Latency Networking
di: Faltelli, Marco, et al.
Pubblicazione: (2024)
di: Faltelli, Marco, et al.
Pubblicazione: (2024)
Multitier Service Migration Framework Based on Mobility Prediction in Mobile Edge Computing
di: Yang, Run, et al.
Pubblicazione: (2024)
di: Yang, Run, et al.
Pubblicazione: (2024)
Contextual Chain: Single-State Ledger Design for Mobile/IoT Networks with Frequent Partitions
di: Kim, Song-Ju
Pubblicazione: (2026)
di: Kim, Song-Ju
Pubblicazione: (2026)
Runtime Verification Containers for Publish/Subscribe Networks
di: Mehran, Ali, et al.
Pubblicazione: (2024)
di: Mehran, Ali, et al.
Pubblicazione: (2024)
Practical Rateless Set Reconciliation
di: Yang, Lei, et al.
Pubblicazione: (2024)
di: Yang, Lei, et al.
Pubblicazione: (2024)
Documenti analoghi
-
XWind: A Cross-site Router for Large Language Model Inference Serving at Renewable Energy Farms
di: Reddy, Tella Rajashekhar, et al.
Pubblicazione: (2026) -
AI Greenferencing: Routing AI Inferencing to Green Modular Data Centers with Heron
di: Reddy, Tella Rajashekhar, et al.
Pubblicazione: (2025) -
Improving training time and GPU utilization in geo-distributed language model training
di: Palak, et al.
Pubblicazione: (2024) -
Local Fast Rerouting with Low Congestion: A Randomized Approach
di: Bankhamer, Gregor, et al.
Pubblicazione: (2020) -
MLTCP: Congestion Control for DNN Training
di: Rajasekaran, Sudarsanan, et al.
Pubblicazione: (2024)