WANSpec: Leveraging Global Compute Capacity for LLM Inference
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Martin, Noah, Dogar, Fahad |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LLMBridge: Reducing Costs to Access LLMs in a Prompt-Centric Internet
von: Martin, Noah, et al.
Veröffentlicht: (2024)
von: Martin, Noah, et al.
Veröffentlicht: (2024)
Towards providing reliable job completion time predictions using PCS
von: Faisal, Abdullah Bin, et al.
Veröffentlicht: (2024)
von: Faisal, Abdullah Bin, et al.
Veröffentlicht: (2024)
inference-fleet-sim: A Queueing-Theory-Grounded Fleet Capacity Planner for LLM Inference
von: Chen, Huamin, et al.
Veröffentlicht: (2026)
von: Chen, Huamin, et al.
Veröffentlicht: (2026)
Distributed On-Device LLM Inference With Over-the-Air Computation
von: Zhang, Kai, et al.
Veröffentlicht: (2025)
von: Zhang, Kai, et al.
Veröffentlicht: (2025)
EdgeShard: Efficient LLM Inference via Collaborative Edge Computing
von: Zhang, Mingjin, et al.
Veröffentlicht: (2024)
von: Zhang, Mingjin, et al.
Veröffentlicht: (2024)
SLICE: SLO-Driven Scheduling for LLM Inference on Edge Computing Devices
von: Chow, Will
Veröffentlicht: (2025)
von: Chow, Will
Veröffentlicht: (2025)
Efficient Routing of Inference Requests across LLM Instances in Cloud-Edge Computing
von: Yu, Shibo, et al.
Veröffentlicht: (2025)
von: Yu, Shibo, et al.
Veröffentlicht: (2025)
Scalable Analysis of Urban Scaling Laws: Leveraging Cloud Computing to Analyze 21,280 Global Cities
von: Li, Zhenhui, et al.
Veröffentlicht: (2024)
von: Li, Zhenhui, et al.
Veröffentlicht: (2024)
More for Less: Integrating Capability-Predominant and Capacity-Predominant Computing
von: Zheng, Zhong, et al.
Veröffentlicht: (2025)
von: Zheng, Zhong, et al.
Veröffentlicht: (2025)
Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads
von: Hu, Cunchen, et al.
Veröffentlicht: (2024)
von: Hu, Cunchen, et al.
Veröffentlicht: (2024)
Federated Inference for Heterogeneous LLM Communication and Collaboration
von: Chen, Zihan, et al.
Veröffentlicht: (2026)
von: Chen, Zihan, et al.
Veröffentlicht: (2026)
Cloud Native System for LLM Inference Serving
von: Xu, Minxian, et al.
Veröffentlicht: (2025)
von: Xu, Minxian, et al.
Veröffentlicht: (2025)
Enabling Dynamic Sparsity in Quantized LLM Inference
von: Wang, Rongxiang, et al.
Veröffentlicht: (2025)
von: Wang, Rongxiang, et al.
Veröffentlicht: (2025)
Understanding the Performance and Power of LLM Inferencing on Edge Accelerators
von: Arya, Mayank, et al.
Veröffentlicht: (2025)
von: Arya, Mayank, et al.
Veröffentlicht: (2025)
From Attention to Disaggregation: Tracing the Evolution of LLM Inference
von: Kumar, Madabattula Rajesh, et al.
Veröffentlicht: (2025)
von: Kumar, Madabattula Rajesh, et al.
Veröffentlicht: (2025)
Towards Resource-Efficient Serverless LLM Inference with SLINFER
von: Xu, Chuhao, et al.
Veröffentlicht: (2025)
von: Xu, Chuhao, et al.
Veröffentlicht: (2025)
Argus: Token Aware Distributed LLM Inference Optimization
von: Wu, Panlong, et al.
Veröffentlicht: (2025)
von: Wu, Panlong, et al.
Veröffentlicht: (2025)
Toward Sustainability-Aware LLM Inference on Edge Clusters
von: Rajashekar, Kolichala, et al.
Veröffentlicht: (2025)
von: Rajashekar, Kolichala, et al.
Veröffentlicht: (2025)
Efficient LLM Inference with Activation Checkpointing and Hybrid Caching
von: Lee, Sanghyeon, et al.
Veröffentlicht: (2025)
von: Lee, Sanghyeon, et al.
Veröffentlicht: (2025)
SYMPHONY: Improving Memory Management for LLM Inference Workloads
von: Agarwal, Saurabh, et al.
Veröffentlicht: (2024)
von: Agarwal, Saurabh, et al.
Veröffentlicht: (2024)
LLM-Emu: Native Runtime Emulation of LLM Inference via Profile-Driven Sampling
von: Da, Wei, et al.
Veröffentlicht: (2026)
von: Da, Wei, et al.
Veröffentlicht: (2026)
AcceLLM: Accelerating LLM Inference using Redundancy for Load Balancing and Data Locality
von: Bournias, Ilias, et al.
Veröffentlicht: (2024)
von: Bournias, Ilias, et al.
Veröffentlicht: (2024)
VQ-LLM: High-performance Code Generation for Vector Quantization Augmented LLM Inference
von: Liu, Zihan, et al.
Veröffentlicht: (2025)
von: Liu, Zihan, et al.
Veröffentlicht: (2025)
Towards Exascale Computing for Astrophysical Simulation Leveraging the Leonardo EuroHPC System
von: Shukla, Nitin, et al.
Veröffentlicht: (2025)
von: Shukla, Nitin, et al.
Veröffentlicht: (2025)
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing
von: Lin, Mao, et al.
Veröffentlicht: (2026)
von: Lin, Mao, et al.
Veröffentlicht: (2026)
Efficient Multi-round LLM Inference over Disaggregated Serving
von: He, Wenhao, et al.
Veröffentlicht: (2026)
von: He, Wenhao, et al.
Veröffentlicht: (2026)
Bandwidth-Aware LLM Inference on Heterogeneous Many-Core Supercomputers
von: Lu, Yao, et al.
Veröffentlicht: (2026)
von: Lu, Yao, et al.
Veröffentlicht: (2026)
Arrow: Adaptive Scheduling Mechanisms for Disaggregated LLM Inference Architecture
von: Wu, Yu, et al.
Veröffentlicht: (2025)
von: Wu, Yu, et al.
Veröffentlicht: (2025)
Amoeba: Runtime Tensor Parallel Transformation for LLM Inference Services
von: Chen, Haoyu, et al.
Veröffentlicht: (2025)
von: Chen, Haoyu, et al.
Veröffentlicht: (2025)
SCOOT: SLO-Oriented Performance Tuning for LLM Inference Engines
von: Cheng, Ke, et al.
Veröffentlicht: (2024)
von: Cheng, Ke, et al.
Veröffentlicht: (2024)
UELLM: A Unified and Efficient Approach for LLM Inference Serving
von: He, Yiyuan, et al.
Veröffentlicht: (2024)
von: He, Yiyuan, et al.
Veröffentlicht: (2024)
Parallax: Efficient LLM Inference Service over Decentralized Environment
von: Tong, Chris, et al.
Veröffentlicht: (2025)
von: Tong, Chris, et al.
Veröffentlicht: (2025)
Decentralized LLM Inference over Edge Networks with Energy Harvesting
von: Khoshsirat, Aria, et al.
Veröffentlicht: (2024)
von: Khoshsirat, Aria, et al.
Veröffentlicht: (2024)
LLM-CoOpt: A Co-Design and Optimization Framework for Efficient LLM Inference on Heterogeneous Platforms
von: Kong, Jie, et al.
Veröffentlicht: (2026)
von: Kong, Jie, et al.
Veröffentlicht: (2026)
Offline Energy-Optimal LLM Serving: Workload-Based Energy Models for LLM Inference on Heterogeneous Systems
von: Wilkins, Grant, et al.
Veröffentlicht: (2024)
von: Wilkins, Grant, et al.
Veröffentlicht: (2024)
Online Optimization of DNN Inference Network Utility in Collaborative Edge Computing
von: Li, Rui, et al.
Veröffentlicht: (2024)
von: Li, Rui, et al.
Veröffentlicht: (2024)
Adaptive Configuration Selection for Multi-Model Inference Pipelines in Edge Computing
von: Sheng, Jinhao, et al.
Veröffentlicht: (2025)
von: Sheng, Jinhao, et al.
Veröffentlicht: (2025)
PipeMax: Enhancing Offline LLM Inference on Commodity GPU Servers
von: Zhang, Hongbin, et al.
Veröffentlicht: (2026)
von: Zhang, Hongbin, et al.
Veröffentlicht: (2026)
Taming Request Imbalance: SLO-Aware Scheduling for Disaggregated LLM Inference
von: Wang, Qipeng
Veröffentlicht: (2026)
von: Wang, Qipeng
Veröffentlicht: (2026)
SiDP: Memory-Efficient Data Parallelism for Offline LLM Inference
von: Zhao, Alan, et al.
Veröffentlicht: (2026)
von: Zhao, Alan, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
LLMBridge: Reducing Costs to Access LLMs in a Prompt-Centric Internet
von: Martin, Noah, et al.
Veröffentlicht: (2024) -
Towards providing reliable job completion time predictions using PCS
von: Faisal, Abdullah Bin, et al.
Veröffentlicht: (2024) -
inference-fleet-sim: A Queueing-Theory-Grounded Fleet Capacity Planner for LLM Inference
von: Chen, Huamin, et al.
Veröffentlicht: (2026) -
Distributed On-Device LLM Inference With Over-the-Air Computation
von: Zhang, Kai, et al.
Veröffentlicht: (2025) -
EdgeShard: Efficient LLM Inference via Collaborative Edge Computing
von: Zhang, Mingjin, et al.
Veröffentlicht: (2024)