Improving the End-to-End Efficiency of Offline Inference for Multi-LLM Applications Based on Sampling and Simulation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Fang, Jingzhi, Shen, Yanyan, Wang, Yue, Chen, Lei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ScaleLLM: A Resource-Frugal LLM Serving Framework by Optimizing End-to-End Efficiency
von: Yao, Yuhang, et al.
Veröffentlicht: (2024)
von: Yao, Yuhang, et al.
Veröffentlicht: (2024)
Deal: Distributed End-to-End GNN Inference for All Nodes
von: Chen, Shiyang, et al.
Veröffentlicht: (2025)
von: Chen, Shiyang, et al.
Veröffentlicht: (2025)
nncase: An End-to-End Compiler for Efficient LLM Deployment on Heterogeneous Storage Architectures
von: Guo, Hui, et al.
Veröffentlicht: (2025)
von: Guo, Hui, et al.
Veröffentlicht: (2025)
StraightLine: An End-to-End Resource-Aware Scheduler for Machine Learning Application Requests
von: Ching, Cheng-Wei, et al.
Veröffentlicht: (2024)
von: Ching, Cheng-Wei, et al.
Veröffentlicht: (2024)
Understanding and Improving Communication Performance in Multi-node LLM Inference
von: Singhania, Prajwal, et al.
Veröffentlicht: (2025)
von: Singhania, Prajwal, et al.
Veröffentlicht: (2025)
An End-to-End DNN Inference Framework for the SpiNNaker2 Neuromorphic MPSoC
von: Jobst, Matthias, et al.
Veröffentlicht: (2025)
von: Jobst, Matthias, et al.
Veröffentlicht: (2025)
Communication Resources Constrained Hierarchical Federated Learning for End-to-End Autonomous Driving
von: Kou, Wei-Bin, et al.
Veröffentlicht: (2023)
von: Kou, Wei-Bin, et al.
Veröffentlicht: (2023)
Staggered Batch Scheduling: Co-optimizing Time-to-First-Token and Throughput for High-Efficiency LLM Inference
von: Tian, Jian, et al.
Veröffentlicht: (2025)
von: Tian, Jian, et al.
Veröffentlicht: (2025)
Glinthawk: A Two-Tiered Architecture for Offline LLM Inference
von: Hamadanian, Pouya, et al.
Veröffentlicht: (2025)
von: Hamadanian, Pouya, et al.
Veröffentlicht: (2025)
FT-Transformer: Resilient and Reliable Transformer with End-to-End Fault Tolerant Attention
von: Dai, Huangliang, et al.
Veröffentlicht: (2025)
von: Dai, Huangliang, et al.
Veröffentlicht: (2025)
MobiZO: Enabling Efficient LLM Fine-Tuning at the Edge via Inference Engines
von: Gao, Lei, et al.
Veröffentlicht: (2024)
von: Gao, Lei, et al.
Veröffentlicht: (2024)
FDC: Fast KV Dimensionality Compression for Efficient LLM Inference
von: Zhang, Zeyu, et al.
Veröffentlicht: (2024)
von: Zhang, Zeyu, et al.
Veröffentlicht: (2024)
PecSched: Preemptive and Efficient Cluster Scheduling for LLM Inference
von: Zhang, Zeyu, et al.
Veröffentlicht: (2024)
von: Zhang, Zeyu, et al.
Veröffentlicht: (2024)
End-to-End Verifiable Decentralized Federated Learning
von: Lee, Chaehyeon, et al.
Veröffentlicht: (2024)
von: Lee, Chaehyeon, et al.
Veröffentlicht: (2024)
Mind the Memory Gap: Unveiling GPU Bottlenecks in Large-Batch LLM Inference
von: Recasens, Pol G., et al.
Veröffentlicht: (2025)
von: Recasens, Pol G., et al.
Veröffentlicht: (2025)
STAR: Decode-Phase Rescheduling for LLM Inference
von: Wang, Zhibin, et al.
Veröffentlicht: (2025)
von: Wang, Zhibin, et al.
Veröffentlicht: (2025)
LLM Inference at the Edge: Mobile, NPU, and GPU Performance Efficiency Trade-offs Under Sustained Load
von: Tummalapalli, Pranay, et al.
Veröffentlicht: (2026)
von: Tummalapalli, Pranay, et al.
Veröffentlicht: (2026)
HyGen: Efficient LLM Serving via Elastic Online-Offline Request Co-location
von: Sun, Ting, et al.
Veröffentlicht: (2025)
von: Sun, Ting, et al.
Veröffentlicht: (2025)
Agglomerative Federated Learning: Empowering Larger Model Training via End-Edge-Cloud Collaboration
von: Wu, Zhiyuan, et al.
Veröffentlicht: (2023)
von: Wu, Zhiyuan, et al.
Veröffentlicht: (2023)
ConServe: Fine-Grained GPU Harvesting for LLM Online and Offline Co-Serving
von: Qiao, Yifan, et al.
Veröffentlicht: (2024)
von: Qiao, Yifan, et al.
Veröffentlicht: (2024)
Beyond Model Scale Limits: End-Edge-Cloud Federated Learning with Self-Rectified Knowledge Agglomeration
von: Wu, Zhiyuan, et al.
Veröffentlicht: (2025)
von: Wu, Zhiyuan, et al.
Veröffentlicht: (2025)
Harnessing Your DRAM and SSD for Sustainable and Accessible LLM Inference with Mixed-Precision and Multi-level Caching
von: Peng, Jie, et al.
Veröffentlicht: (2024)
von: Peng, Jie, et al.
Veröffentlicht: (2024)
Collaborative Speculative Inference for Efficient LLM Inference Serving
von: Gao, Luyao, et al.
Veröffentlicht: (2025)
von: Gao, Luyao, et al.
Veröffentlicht: (2025)
Salted Inference: Enhancing Privacy while Maintaining Efficiency of Split Inference in Mobile Computing
von: Malekzadeh, Mohammad, et al.
Veröffentlicht: (2023)
von: Malekzadeh, Mohammad, et al.
Veröffentlicht: (2023)
CALVO: Improve Serving Efficiency for LLM Inferences with Intense Network Demands
von: Wang, Weiye, et al.
Veröffentlicht: (2026)
von: Wang, Weiye, et al.
Veröffentlicht: (2026)
HACK: Homomorphic Acceleration via Compression of the Key-Value Cache for Disaggregated LLM Inference
von: Zhang, Zeyu, et al.
Veröffentlicht: (2025)
von: Zhang, Zeyu, et al.
Veröffentlicht: (2025)
MARLIN: Multi-Agent Game-Theoretic Reinforcement Learning for Sustainable LLM Inference in Cloud Datacenters
von: Moore, H., et al.
Veröffentlicht: (2026)
von: Moore, H., et al.
Veröffentlicht: (2026)
MultiTASC++: A Continuously Adaptive Scheduler for Edge-Based Multi-Device Cascade Inference
von: Nikolaidis, Sokratis, et al.
Veröffentlicht: (2024)
von: Nikolaidis, Sokratis, et al.
Veröffentlicht: (2024)
Energy Use of AI Inference: Efficiency Pathways and Test-Time Compute
von: Oviedo, Felipe, et al.
Veröffentlicht: (2025)
von: Oviedo, Felipe, et al.
Veröffentlicht: (2025)
Pie: Pooling CPU Memory for LLM Inference
von: Xu, Yi, et al.
Veröffentlicht: (2024)
von: Xu, Yi, et al.
Veröffentlicht: (2024)
The Workload-Router-Pool Architecture for LLM Inference Optimization: A Vision Paper from the vLLM Semantic Router Project
von: Chen, Huamin, et al.
Veröffentlicht: (2026)
von: Chen, Huamin, et al.
Veröffentlicht: (2026)
Local-Cloud Inference Offloading for LLMs in Multi-Modal, Multi-Task, Multi-Dialogue Settings
von: Yuan, Liangqi, et al.
Veröffentlicht: (2025)
von: Yuan, Liangqi, et al.
Veröffentlicht: (2025)
Beyond End-to-End: Dynamic Chain Optimization for Private LLM Adaptation on the Edge
von: Wu, Yebo, et al.
Veröffentlicht: (2026)
von: Wu, Yebo, et al.
Veröffentlicht: (2026)
Optimal Scheduling Algorithms for LLM Inference: Theory and Practice
von: Bari, Agrim, et al.
Veröffentlicht: (2025)
von: Bari, Agrim, et al.
Veröffentlicht: (2025)
Improving the Efficiency of a Deep Reinforcement Learning-Based Power Management System for HPC Clusters Using Curriculum Learning
von: Budiarjo, Thomas, et al.
Veröffentlicht: (2025)
von: Budiarjo, Thomas, et al.
Veröffentlicht: (2025)
No Request Left Behind: Tackling Heterogeneity in Long-Context LLM Inference with Medha
von: Agrawal, Amey, et al.
Veröffentlicht: (2024)
von: Agrawal, Amey, et al.
Veröffentlicht: (2024)
Canvas: End-to-End Kernel Architecture Search in Neural Networks
von: Zhao, Chenggang, et al.
Veröffentlicht: (2023)
von: Zhao, Chenggang, et al.
Veröffentlicht: (2023)
A Survey of End-to-End Modeling for Distributed DNN Training: Workloads, Simulators, and TCO
von: Svedas, Jonas, et al.
Veröffentlicht: (2025)
von: Svedas, Jonas, et al.
Veröffentlicht: (2025)
Exploring Influence Factors on LLM Suitability for No-Code Development of End User IoT Applications
von: Wang, Minghe, et al.
Veröffentlicht: (2025)
von: Wang, Minghe, et al.
Veröffentlicht: (2025)
eMoE: Task-aware Memory Efficient Mixture-of-Experts-Based (MoE) Model Inference
von: Tairin, Suraiya, et al.
Veröffentlicht: (2025)
von: Tairin, Suraiya, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
ScaleLLM: A Resource-Frugal LLM Serving Framework by Optimizing End-to-End Efficiency
von: Yao, Yuhang, et al.
Veröffentlicht: (2024) -
Deal: Distributed End-to-End GNN Inference for All Nodes
von: Chen, Shiyang, et al.
Veröffentlicht: (2025) -
nncase: An End-to-End Compiler for Efficient LLM Deployment on Heterogeneous Storage Architectures
von: Guo, Hui, et al.
Veröffentlicht: (2025) -
StraightLine: An End-to-End Resource-Aware Scheduler for Machine Learning Application Requests
von: Ching, Cheng-Wei, et al.
Veröffentlicht: (2024) -
Understanding and Improving Communication Performance in Multi-node LLM Inference
von: Singhania, Prajwal, et al.
Veröffentlicht: (2025)