MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services
Fuente:
arXiv
Saved in:
| Main Authors: | Tang, Lingfeng, Zhang, Daoping, Chen, Junjie, Huang, Peihao, Jin, Feng, Xu, Chengguang, Chen, Yuxin, Sun, Feiqiang, Chen, Guo |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
On the Bandwidth Consumption of Blockchains
by: Lebedev, Andrei, et al.
Published: (2026)
by: Lebedev, Andrei, et al.
Published: (2026)
TraDE: Network and Traffic-aware Adaptive Scheduling for Microservices Under Dynamics
by: Chen, Ming, et al.
Published: (2024)
by: Chen, Ming, et al.
Published: (2024)
JANUS: Resilient and Adaptive Data Transmission for Enabling Timely and Efficient Cross-Facility Scientific Workflows
by: Esaulov, Vladislav, et al.
Published: (2025)
by: Esaulov, Vladislav, et al.
Published: (2025)
Accelerator-as-a-Service in Public Clouds: An Intra-Host Traffic Management View for Performance Isolation in the Wild
by: Zhao, Jiechen, et al.
Published: (2024)
by: Zhao, Jiechen, et al.
Published: (2024)
Local Rendezvous Hashing: Bounded Loads and Minimal Churn via Cache-Local Candidates
by: Guan, Yongjie
Published: (2025)
by: Guan, Yongjie
Published: (2025)
EvalNet: A Practical Toolchain for Generation and Analysis of Extreme-Scale Interconnects
by: Besta, Maciej, et al.
Published: (2021)
by: Besta, Maciej, et al.
Published: (2021)
Performance and Stability of Barrier Mode Parallel Systems with Heterogeneous and Redundant Jobs
by: Walker, Brenton, et al.
Published: (2025)
by: Walker, Brenton, et al.
Published: (2025)
PANDAS: Peer-to-peer, Adaptive Networking for Data Availability Sampling within Ethereum Consensus Timebounds
by: Pigaglio, Matthieu, et al.
Published: (2025)
by: Pigaglio, Matthieu, et al.
Published: (2025)
Joint Network Slicing, Routing, and In-Network Computing for Energy-Efficient 6G
by: Sasan, Zeinab, et al.
Published: (2024)
by: Sasan, Zeinab, et al.
Published: (2024)
FPsPIN: An FPGA-based Open-Hardware Research Platform for Processing in the Network
by: Schneider, Timo, et al.
Published: (2024)
by: Schneider, Timo, et al.
Published: (2024)
Energy-Efficient and High-Performance Data Transfers with DRL Agents
by: Jamil, Hasibul, et al.
Published: (2025)
by: Jamil, Hasibul, et al.
Published: (2025)
CASPER: Carbon-Aware Scheduling and Provisioning for Distributed Web Services
by: Souza, Abel, et al.
Published: (2024)
by: Souza, Abel, et al.
Published: (2024)
Proactive Service Assurance in 5G and B5G Networks: A Closed-Loop Algorithm for End-to-End Network Slicing
by: Tran, Nguyen Phuc, et al.
Published: (2024)
by: Tran, Nguyen Phuc, et al.
Published: (2024)
AFLL: Real-time Load Stabilization for MMO Game Servers Based on Circular Causality Learning
by: Kang, Shinsuk, et al.
Published: (2026)
by: Kang, Shinsuk, et al.
Published: (2026)
GigaAPI for GPU Parallelization
by: Suvarna, M., et al.
Published: (2025)
by: Suvarna, M., et al.
Published: (2025)
Parallelizing a modern GPU simulator
by: Huerta, Rodrigo, et al.
Published: (2025)
by: Huerta, Rodrigo, et al.
Published: (2025)
Edge-First Language Model Inference: Models, Metrics, and Tradeoffs
by: Jang, SiYoung, et al.
Published: (2025)
by: Jang, SiYoung, et al.
Published: (2025)
LLAMP: Assessing Network Latency Tolerance of HPC Applications with Linear Programming
by: Shen, Siyuan, et al.
Published: (2024)
by: Shen, Siyuan, et al.
Published: (2024)
On Capacity and Delay of Wireless Networks with Node Failures
by: Li, Wei, et al.
Published: (2026)
by: Li, Wei, et al.
Published: (2026)
A Case for CATS: A Conductor-driven Asymmetric Transport Scheme for Semantic Prioritization
by: Rizvi, Syed Muhammad Aqdas
Published: (2026)
by: Rizvi, Syed Muhammad Aqdas
Published: (2026)
Network-Aware Scheduling for Remote Gate Execution in Quantum Data Centers
by: Pouryousef, Shahrooz, et al.
Published: (2025)
by: Pouryousef, Shahrooz, et al.
Published: (2025)
Benchmarking Quantum Data Center Architectures: A Performance and Scalability Perspective
by: Pouryousef, Shahrooz, et al.
Published: (2026)
by: Pouryousef, Shahrooz, et al.
Published: (2026)
QECO: A QoE-Oriented Computation Offloading Algorithm based on Deep Reinforcement Learning for Mobile Edge Computing
by: Rahmaty, Iman, et al.
Published: (2023)
by: Rahmaty, Iman, et al.
Published: (2023)
Generative AI on the Edge: Architecture and Performance Evaluation
by: Nezami, Zeinab, et al.
Published: (2024)
by: Nezami, Zeinab, et al.
Published: (2024)
From Skew to Symmetry: Node-Interconnect Multi-Path Balancing with Execution-time Planning for Modern GPU Clusters
by: Yao, Jinghan, et al.
Published: (2026)
by: Yao, Jinghan, et al.
Published: (2026)
Ultra Ethernet's Design Principles and Architectural Innovations
by: Hoefler, Torsten, et al.
Published: (2025)
by: Hoefler, Torsten, et al.
Published: (2025)
Temporal-Aware GPU Resource Allocation for Distributed LLM Inference via Reinforcement Learning
by: Du, Chengze, et al.
Published: (2025)
by: Du, Chengze, et al.
Published: (2025)
Compiler Support for Speculation in Decoupled Access/Execute Architectures
by: Szafarczyk, Robert, et al.
Published: (2025)
by: Szafarczyk, Robert, et al.
Published: (2025)
Optimizing Intra-Container Communication with Memory Protection Keys: A Novel Approach to Secure and Efficient Microservice Interaction
by: Yashu, Fnu, et al.
Published: (2025)
by: Yashu, Fnu, et al.
Published: (2025)
Characterizing and Optimizing LLM Inference Workloads on CPU-GPU Coupled Architectures
by: Vellaisamy, Prabhu, et al.
Published: (2025)
by: Vellaisamy, Prabhu, et al.
Published: (2025)
Multi-stage Flow Scheduling for LLM Serving
by: Sun, Yijun, et al.
Published: (2026)
by: Sun, Yijun, et al.
Published: (2026)
PerLLM: Personalized Inference Scheduling with Edge-Cloud Collaboration for Diverse LLM Services
by: Yang, Zheming, et al.
Published: (2024)
by: Yang, Zheming, et al.
Published: (2024)
FAST: An Efficient Scheduler for All-to-All GPU Communication
by: Lei, Yiran, et al.
Published: (2025)
by: Lei, Yiran, et al.
Published: (2025)
InfiniteHBD: Building Datacenter-Scale High-Bandwidth Domain for LLM with Optical Circuit Switching Transceivers
by: Shou, Chenchen, et al.
Published: (2025)
by: Shou, Chenchen, et al.
Published: (2025)
PeerSync: Accelerating Containerized Service Delivery at the Network Edge
by: Deng, Yinuo, et al.
Published: (2025)
by: Deng, Yinuo, et al.
Published: (2025)
Dynamic Hierarchical Birkhoff-von Neumann Decomposition for All-to-All GPU Communication
by: Wu, Yen-Chieh, et al.
Published: (2026)
by: Wu, Yen-Chieh, et al.
Published: (2026)
An Online Fragmentation-Aware GPU Scheduler for Multi-Tenant MIG-based Clouds
by: Zambianco, Marco, et al.
Published: (2025)
by: Zambianco, Marco, et al.
Published: (2025)
Can Asymmetric Tile Buffering Be Beneficial?
by: Wang, Chengyue, et al.
Published: (2025)
by: Wang, Chengyue, et al.
Published: (2025)
DualPath: Breaking the Storage Bandwidth Bottleneck in Agentic LLM Inference
by: Wu, Yongtong, et al.
Published: (2026)
by: Wu, Yongtong, et al.
Published: (2026)
ZipServ: Fast and Memory-Efficient LLM Inference with Hardware-Aware Lossless Compression
by: Fan, Ruibo, et al.
Published: (2026)
by: Fan, Ruibo, et al.
Published: (2026)
Similar Items
-
On the Bandwidth Consumption of Blockchains
by: Lebedev, Andrei, et al.
Published: (2026) -
TraDE: Network and Traffic-aware Adaptive Scheduling for Microservices Under Dynamics
by: Chen, Ming, et al.
Published: (2024) -
JANUS: Resilient and Adaptive Data Transmission for Enabling Timely and Efficient Cross-Facility Scientific Workflows
by: Esaulov, Vladislav, et al.
Published: (2025) -
Accelerator-as-a-Service in Public Clouds: An Intra-Host Traffic Management View for Performance Isolation in the Wild
by: Zhao, Jiechen, et al.
Published: (2024) -
Local Rendezvous Hashing: Bounded Loads and Minimal Churn via Cache-Local Candidates
by: Guan, Yongjie
Published: (2025)