Network and Systems Performance Characterization of MCP-Enabled LLM Agents
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Ding, Zihao, Zhu, Mufeng, Liu, Yao |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Resilient AI Supercomputer Networking using MRC and SRv6
par: Araujo, Joao, et autres
Publié: (2026)
par: Araujo, Joao, et autres
Publié: (2026)
AIvailable: A Software-Defined Architecture for LLM-as-a-Service on Heterogeneous and Legacy GPUs
par: Antunes, Pedro, et autres
Publié: (2025)
par: Antunes, Pedro, et autres
Publié: (2025)
Architecture-Aware LLM Inference Optimization on AMD Instinct GPUs: A Comprehensive Benchmark and Deployment Study
par: Georgiou, Athos
Publié: (2026)
par: Georgiou, Athos
Publié: (2026)
Towards Message Brokers for Generative AI: Survey, Challenges, and Opportunities
par: Saleh, Alaa, et autres
Publié: (2023)
par: Saleh, Alaa, et autres
Publié: (2023)
FedMon: Federated eBPF Monitoring for Distributed Anomaly Detection in Multi-Cluster Cloud Environments
par: Zehra, Sehar, et autres
Publié: (2025)
par: Zehra, Sehar, et autres
Publié: (2025)
Towards Policy-Enabled Multi-Hop Routing for Cross-Chain Message Delivery
par: Rezaei, Amin, et autres
Publié: (2026)
par: Rezaei, Amin, et autres
Publié: (2026)
Rotary GPU: Exploring Local Execution Paths for Large Mixture-of-Experts Models Under Limited GPU Memory
par: Jo, Myeong Jun
Publié: (2026)
par: Jo, Myeong Jun
Publié: (2026)
Neural Router: Semantic Content Matching for Agentic AI
par: Lovén, Lauri, et autres
Publié: (2026)
par: Lovén, Lauri, et autres
Publié: (2026)
CooperLLM: Cloud-Edge-End Cooperative Federated Fine-tuning for LLMs via ZOO-based Gradient Correction
par: Sun, He, et autres
Publié: (2026)
par: Sun, He, et autres
Publié: (2026)
Service Discovery-Based Hybrid Network Middleware for Efficient Communication in Distributed Robotic Systems
par: Sang, Shiyao, et autres
Publié: (2025)
par: Sang, Shiyao, et autres
Publié: (2025)
AAFLOW: Scalable Patterns for Agentic AI Workflows
par: Sarker, Arup Kumar, et autres
Publié: (2026)
par: Sarker, Arup Kumar, et autres
Publié: (2026)
Comparative Analysis of Large Language Model Inference Serving Systems: A Performance Study of vLLM and HuggingFace TGI
par: Kolluru, Saicharan
Publié: (2025)
par: Kolluru, Saicharan
Publié: (2025)
Swing: Short-cutting Rings for Higher Bandwidth Allreduce
par: De Sensi, Daniele, et autres
Publié: (2024)
par: De Sensi, Daniele, et autres
Publié: (2024)
Joint Task Offloading and Routing in Wireless Multi-hop Networks Using Biased Backpressure Algorithm
par: Zhao, Zhongyuan, et autres
Publié: (2024)
par: Zhao, Zhongyuan, et autres
Publié: (2024)
POD-Attention: Unlocking Full Prefill-Decode Overlap for Faster LLM Inference
par: Kamath, Aditya K, et autres
Publié: (2024)
par: Kamath, Aditya K, et autres
Publié: (2024)
Predictive Multi-Tier Memory Management for KV Cache in Large-Scale GPU Inference
par: Ganjihal, Sanjeev Rao
Publié: (2026)
par: Ganjihal, Sanjeev Rao
Publié: (2026)
Training LLMs on HPC Systems: Best Practices from the OpenGPT-X Project
par: Penke, Carolin, et autres
Publié: (2025)
par: Penke, Carolin, et autres
Publié: (2025)
Parameter-Efficient and Personalized Federated Training of Generative Models at the Edge
par: Khan, Kabir, et autres
Publié: (2025)
par: Khan, Kabir, et autres
Publié: (2025)
SLA Management in Reconfigurable Multi-Agent RAG: A Systems Approach to Question Answering
par: Iannelli, Michael, et autres
Publié: (2024)
par: Iannelli, Michael, et autres
Publié: (2024)
Collective Communication for 100k+ GPUs
par: Si, Min, et autres
Publié: (2025)
par: Si, Min, et autres
Publié: (2025)
Directives for Function Offloading in 5G Networks Based on a Performance Characteristics Analysis
par: Dettinger, Falk, et autres
Publié: (2025)
par: Dettinger, Falk, et autres
Publié: (2025)
Asynchronous Pipeline Parallelism for Real-Time Multilingual Lip Synchronization in Video Communication Systems
par: Caglar, Eren, et autres
Publié: (2025)
par: Caglar, Eren, et autres
Publié: (2025)
Cognitive Infrastructure: A Unified DCIM Framework for AI Data Centers
par: Sunkara, Krishna Chaitanya
Publié: (2026)
par: Sunkara, Krishna Chaitanya
Publié: (2026)
Konnektor: Connection Protocol for Ensuring Peer Uniqueness in Decentralized P2P Networks
par: Ozkan, Onur
Publié: (2024)
par: Ozkan, Onur
Publié: (2024)
Autonomous Trajectory Optimization for UAVs in Disaster Zone Using Henry Gas Optimization Scheme
par: Qadir, Zakria, et autres
Publié: (2025)
par: Qadir, Zakria, et autres
Publié: (2025)
Augmenting the FedProx Algorithm by Minimizing Convergence
par: Sarkar, Anomitra, et autres
Publié: (2024)
par: Sarkar, Anomitra, et autres
Publié: (2024)
DSDE: Dynamic Speculative Decoding with KLD Stability for Real-World Serving
par: Yang, Mingyu, et autres
Publié: (2025)
par: Yang, Mingyu, et autres
Publié: (2025)
DynamiQ: Accelerating Gradient Synchronization using Compressed Multi-hop All-reduce
par: Han, Wenchen, et autres
Publié: (2026)
par: Han, Wenchen, et autres
Publié: (2026)
A Survey on Heterogeneous Computing Using SmartNICs and Emerging Data Processing Units
par: Tibbetts, Nathan, et autres
Publié: (2025)
par: Tibbetts, Nathan, et autres
Publié: (2025)
Moonshot: Optimizing Chain-Based Rotating Leader BFT via Optimistic Proposals
par: Doidge, Isaac, et autres
Publié: (2024)
par: Doidge, Isaac, et autres
Publié: (2024)
Toward Architecture-Aware Evaluation Metrics for LLM Agents
par: Souza, Débora, et autres
Publié: (2026)
par: Souza, Débora, et autres
Publié: (2026)
ABACUS: A FinOps Service for Cloud Cost Optimization
par: Deochake, Saurabh
Publié: (2024)
par: Deochake, Saurabh
Publié: (2024)
Benchmarking Federated Learning for Throughput Prediction in 5G Live Streaming Applications
par: Dutta, Yuvraj, et autres
Publié: (2025)
par: Dutta, Yuvraj, et autres
Publié: (2025)
ATTNChecker: Highly-Optimized Fault Tolerant Attention for Large Language Model Training
par: Liang, Yuhang, et autres
Publié: (2024)
par: Liang, Yuhang, et autres
Publié: (2024)
Nezha: Deployable and High-Performance Consensus Using Synchronized Clocks
par: Geng, Jinkun, et autres
Publié: (2022)
par: Geng, Jinkun, et autres
Publié: (2022)
Exploring Micro Frontends: A Case Study Application in E-Commerce
par: Kojo, Ricardo Hideki Hangai, et autres
Publié: (2025)
par: Kojo, Ricardo Hideki Hangai, et autres
Publié: (2025)
LLM-Driven Adaptive 6G-Ready Wireless Body Area Networks: Survey and Framework
par: Torkamani, Mohammad Jalili, et autres
Publié: (2025)
par: Torkamani, Mohammad Jalili, et autres
Publié: (2025)
ConfigSpec: Profiling-Based Configuration Selection for Distributed Edge--Cloud Speculative LLM Serving
par: Li, Xiangchen, et autres
Publié: (2026)
par: Li, Xiangchen, et autres
Publié: (2026)
Hive: A Multi-Agent Infrastructure for Algorithm- and Task-Level Scaling
par: Luo, Zizhang, et autres
Publié: (2026)
par: Luo, Zizhang, et autres
Publié: (2026)
Flash-Fusion: Enabling Expressive, Low-Latency Queries on IoT Sensor Streams with LLMs
par: Patherya, Kausar, et autres
Publié: (2025)
par: Patherya, Kausar, et autres
Publié: (2025)
Documents similaires
-
Resilient AI Supercomputer Networking using MRC and SRv6
par: Araujo, Joao, et autres
Publié: (2026) -
AIvailable: A Software-Defined Architecture for LLM-as-a-Service on Heterogeneous and Legacy GPUs
par: Antunes, Pedro, et autres
Publié: (2025) -
Architecture-Aware LLM Inference Optimization on AMD Instinct GPUs: A Comprehensive Benchmark and Deployment Study
par: Georgiou, Athos
Publié: (2026) -
Towards Message Brokers for Generative AI: Survey, Challenges, and Opportunities
par: Saleh, Alaa, et autres
Publié: (2023) -
FedMon: Federated eBPF Monitoring for Distributed Anomaly Detection in Multi-Cluster Cloud Environments
par: Zehra, Sehar, et autres
Publié: (2025)