DualPath: Breaking the Storage Bandwidth Bottleneck in Agentic LLM Inference
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Wu, Yongtong, Chen, Shaoyuan, Zhong, Yinmin, Huang, Rilin, Tan, Yixuan, Zhang, Wentao, Zhang, Liyue, Zhou, Shangyan, Liu, Yuxuan, Zhou, Shunfeng, Zhang, Mingxing, Jin, Xin, Huang, Panpan |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services
par: Tang, Lingfeng, et autres
Publié: (2025)
par: Tang, Lingfeng, et autres
Publié: (2025)
Insights into DeepSeek-V3: Scaling Challenges and Reflections on Hardware for AI Architectures
par: Zhao, Chenggang, et autres
Publié: (2025)
par: Zhao, Chenggang, et autres
Publié: (2025)
Breaking the Aggregation Bottleneck in Federated Recommendation: A Personalized Model Merging Approach
par: Chen, Jundong, et autres
Publié: (2025)
par: Chen, Jundong, et autres
Publié: (2025)
Breaking the Storage-Bandwidth Tradeoff in Distributed Storage with Quantum Entanglement
par: Hu, Lei, et autres
Publié: (2026)
par: Hu, Lei, et autres
Publié: (2026)
Learning the Optimal Path and DNN Partition for Collaborative Edge Inference
par: Huang, Yin, et autres
Publié: (2024)
par: Huang, Yin, et autres
Publié: (2024)
Fast Distributed Inference Serving for Large Language Models
par: Wu, Bingyang, et autres
Publié: (2023)
par: Wu, Bingyang, et autres
Publié: (2023)
Seer: Proactive Revenue-Aware Scheduling for Live Streaming Services in Crowdsourced Cloud-Edge Platforms
par: Huang, Shaoyuan, et autres
Publié: (2024)
par: Huang, Shaoyuan, et autres
Publié: (2024)
TCDM Burst Access: Breaking the Bandwidth Barrier in Shared-L1 RVV Clusters Beyond 1000 FPUs
par: Shen, Diyou, et autres
Publié: (2025)
par: Shen, Diyou, et autres
Publié: (2025)
Efficient Heterogeneous Large Language Model Decoding with Model-Attention Disaggregation
par: Chen, Shaoyuan, et autres
Publié: (2024)
par: Chen, Shaoyuan, et autres
Publié: (2024)
WOC: Dual-Path Weighted Object Consensus Made Efficient
par: Fonseca, Tanisha, et autres
Publié: (2025)
par: Fonseca, Tanisha, et autres
Publié: (2025)
AdaOper: Energy-efficient and Responsive Concurrent DNN Inference on Mobile Devices
par: Lin, Zheng, et autres
Publié: (2024)
par: Lin, Zheng, et autres
Publié: (2024)
ResBM: Residual Bottleneck Models for Low-Bandwidth Pipeline Parallelism
par: Aboudib, Alan, et autres
Publié: (2026)
par: Aboudib, Alan, et autres
Publié: (2026)
Fire-Flyer AI-HPC: A Cost-Effective Software-Hardware Co-Design for Deep Learning
par: An, Wei, et autres
Publié: (2024)
par: An, Wei, et autres
Publié: (2024)
TokenLake: A Unified Segment-level Prefix Cache Pool for Fine-grained Elastic Long-Context LLM Serving
par: Wu, Bingyang, et autres
Publié: (2025)
par: Wu, Bingyang, et autres
Publié: (2025)
Bandwidth-Aware LLM Inference on Heterogeneous Many-Core Supercomputers
par: Lu, Yao, et autres
Publié: (2026)
par: Lu, Yao, et autres
Publié: (2026)
Breaking the Capacity Bottleneck in Model-Heterogeneous Federated Learning via Gradual Model Restoration
par: Ma, Chengjie, et autres
Publié: (2025)
par: Ma, Chengjie, et autres
Publié: (2025)
Bandwidth-Aware Network Topology Optimization for Decentralized Learning
par: Shen, Yipeng, et autres
Publié: (2025)
par: Shen, Yipeng, et autres
Publié: (2025)
Bandwidth-Aware and Overlap-Weighted Compression for Communication-Efficient Federated Learning
par: Tang, Zichen, et autres
Publié: (2024)
par: Tang, Zichen, et autres
Publié: (2024)
Challenging GPU Dominance: When CPUs Outperform for On-Device LLM Inference
par: Zhang, Haolin, et autres
Publié: (2025)
par: Zhang, Haolin, et autres
Publié: (2025)
Bandwidth-Aware and Cost-Efficient Pipeline Parallel Scheduling in Geo-Distributed LLM Training
par: Zhang, Han, et autres
Publié: (2026)
par: Zhang, Han, et autres
Publié: (2026)
Surviving Partial Rank Failures in Wide Expert-Parallel MoE Inference
par: Sun, Xun, et autres
Publié: (2026)
par: Sun, Xun, et autres
Publié: (2026)
Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles
par: Arif, Moiz, et autres
Publié: (2026)
par: Arif, Moiz, et autres
Publié: (2026)
Remoe: Towards Efficient and Low-Cost MoE Inference in Serverless Computing
par: Liu, Wentao, et autres
Publié: (2025)
par: Liu, Wentao, et autres
Publié: (2025)
Understanding Bottlenecks for Efficiently Serving LLM Inference With KV Offloading
par: Meng, William, et autres
Publié: (2025)
par: Meng, William, et autres
Publié: (2025)
DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving
par: Zhong, Yinmin, et autres
Publié: (2024)
par: Zhong, Yinmin, et autres
Publié: (2024)
CONCUR: High-Throughput Agentic Batch Inference of LLM via Congestion-Based Concurrency Control
par: Chen, Qiaoling, et autres
Publié: (2026)
par: Chen, Qiaoling, et autres
Publié: (2026)
Haina Storage: A Decentralized Secure Storage Framework Based on Improved Blockchain Structure
par: Zhou, Zijian, et autres
Publié: (2024)
par: Zhou, Zijian, et autres
Publié: (2024)
AdapTBF: Decentralized Bandwidth Control via Adaptive Token Borrowing for HPC Storage
par: Rashid, Md Hasanur, et autres
Publié: (2026)
par: Rashid, Md Hasanur, et autres
Publié: (2026)
Collaborative Speculative Inference for Efficient LLM Inference Serving
par: Gao, Luyao, et autres
Publié: (2025)
par: Gao, Luyao, et autres
Publié: (2025)
A 1024 RV-Cores Shared-L1 Cluster with High Bandwidth Memory Link for Low-Latency 6G-SDR
par: Zhang, Yichao, et autres
Publié: (2024)
par: Zhang, Yichao, et autres
Publié: (2024)
Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism
par: Wei, Jinhui, et autres
Publié: (2025)
par: Wei, Jinhui, et autres
Publié: (2025)
OnePiece: A Large-Scale Distributed Inference System with RDMA for Complex AI-Generated Content (AIGC) Workflows
par: Chen, June, et autres
Publié: (2026)
par: Chen, June, et autres
Publié: (2026)
DuoServe-MoE: Dual-Phase Expert Prefetch and Caching for LLM Inference QoS Assurance
par: Zhang, Yuning, et autres
Publié: (2025)
par: Zhang, Yuning, et autres
Publié: (2025)
Distributed Inference Performance Optimization for LLMs on CPUs
par: He, Pujiang, et autres
Publié: (2024)
par: He, Pujiang, et autres
Publié: (2024)
DistTrain: Addressing Model and Data Heterogeneity with Disaggregated Training for Multimodal Large Language Models
par: Zhang, Zili, et autres
Publié: (2024)
par: Zhang, Zili, et autres
Publié: (2024)
VQ-LLM: High-performance Code Generation for Vector Quantization Augmented LLM Inference
par: Liu, Zihan, et autres
Publié: (2025)
par: Liu, Zihan, et autres
Publié: (2025)
Mind the Memory Gap: Unveiling GPU Bottlenecks in Large-Batch LLM Inference
par: Recasens, Pol G., et autres
Publié: (2025)
par: Recasens, Pol G., et autres
Publié: (2025)
KV Cache Compression for Inference Efficiency in LLMs: A Review
par: Liu, Yanyu, et autres
Publié: (2025)
par: Liu, Yanyu, et autres
Publié: (2025)
GriNNder: Breaking the Memory Capacity Wall in Full-Graph GNN Training with Storage Offloading
par: Song, Jaeyong, et autres
Publié: (2026)
par: Song, Jaeyong, et autres
Publié: (2026)
SpecFed: Accelerating Federated LLM Inference with Speculative Decoding and Compressed Transmission
par: Zheng, Ce, et autres
Publié: (2026)
par: Zheng, Ce, et autres
Publié: (2026)
Documents similaires
-
MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services
par: Tang, Lingfeng, et autres
Publié: (2025) -
Insights into DeepSeek-V3: Scaling Challenges and Reflections on Hardware for AI Architectures
par: Zhao, Chenggang, et autres
Publié: (2025) -
Breaking the Aggregation Bottleneck in Federated Recommendation: A Personalized Model Merging Approach
par: Chen, Jundong, et autres
Publié: (2025) -
Breaking the Storage-Bandwidth Tradeoff in Distributed Storage with Quantum Entanglement
par: Hu, Lei, et autres
Publié: (2026) -
Learning the Optimal Path and DNN Partition for Collaborative Edge Inference
par: Huang, Yin, et autres
Publié: (2024)