Etalon: Holistic Performance Evaluation Framework for LLM Inference Systems
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Agrawal, Amey, Agarwal, Anmol, Kedia, Nitin, Mohan, Jayashree, Kundu, Souvik, Kwatra, Nipun, Ramjee, Ramachandran, Tumanov, Alexey |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
On Evaluating Performance of LLM Inference Serving Systems
von: Agrawal, Amey, et al.
Veröffentlicht: (2025)
von: Agrawal, Amey, et al.
Veröffentlicht: (2025)
Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve
von: Agrawal, Amey, et al.
Veröffentlicht: (2024)
von: Agrawal, Amey, et al.
Veröffentlicht: (2024)
TokenWeave: Efficient Compute-Communication Overlap for Distributed LLM Inference
von: Gond, Raja, et al.
Veröffentlicht: (2025)
von: Gond, Raja, et al.
Veröffentlicht: (2025)
Niyama : Breaking the Silos of LLM Inference Serving
von: Goel, Kanishk, et al.
Veröffentlicht: (2025)
von: Goel, Kanishk, et al.
Veröffentlicht: (2025)
Kascade: A Practical Sparse Attention Method for Long-Context LLM Inference
von: Deshmukh, Dhruv, et al.
Veröffentlicht: (2025)
von: Deshmukh, Dhruv, et al.
Veröffentlicht: (2025)
No Request Left Behind: Tackling Heterogeneity in Long-Context LLM Inference with Medha
von: Agrawal, Amey, et al.
Veröffentlicht: (2024)
von: Agrawal, Amey, et al.
Veröffentlicht: (2024)
Sutradhara: An Intelligent Orchestrator-Engine Co-design for Tool-based Agentic Inference
von: Biswas, Anish, et al.
Veröffentlicht: (2026)
von: Biswas, Anish, et al.
Veröffentlicht: (2026)
Vidur: A Large-Scale Simulation Framework For LLM Inference
von: Agrawal, Amey, et al.
Veröffentlicht: (2024)
von: Agrawal, Amey, et al.
Veröffentlicht: (2024)
POD-Attention: Unlocking Full Prefill-Decode Overlap for Faster LLM Inference
von: Kamath, Aditya K, et al.
Veröffentlicht: (2024)
von: Kamath, Aditya K, et al.
Veröffentlicht: (2024)
LLM-42: Enabling Determinism in LLM Inference with Verified Speculation
von: Gond, Raja, et al.
Veröffentlicht: (2026)
von: Gond, Raja, et al.
Veröffentlicht: (2026)
Revati: Transparent GPU-Free Time-Warp Emulation for LLM Serving
von: Agrawal, Amey, et al.
Veröffentlicht: (2026)
von: Agrawal, Amey, et al.
Veröffentlicht: (2026)
SuperFedNAS: Cost-Efficient Federated Neural Architecture Search for On-Device Inference
von: Khare, Alind, et al.
Veröffentlicht: (2023)
von: Khare, Alind, et al.
Veröffentlicht: (2023)
Maya: Optimizing Deep Learning Training Workloads using GPU Runtime Emulation
von: Yarlagadda, Srihas, et al.
Veröffentlicht: (2025)
von: Yarlagadda, Srihas, et al.
Veröffentlicht: (2025)
BlockRaFT: A Distributed Framework for Fault-Tolerant and Scalable Blockchain Nodes
von: Piduguralla, Manaswini, et al.
Veröffentlicht: (2026)
von: Piduguralla, Manaswini, et al.
Veröffentlicht: (2026)
ModServe: Modality- and Stage-Aware Resource Disaggregation for Scalable Multimodal Model Serving
von: Qiu, Haoran, et al.
Veröffentlicht: (2025)
von: Qiu, Haoran, et al.
Veröffentlicht: (2025)
SYMPHONY: Improving Memory Management for LLM Inference Workloads
von: Agarwal, Saurabh, et al.
Veröffentlicht: (2024)
von: Agarwal, Saurabh, et al.
Veröffentlicht: (2024)
Agentic AI Workload Characteristics
von: Yuan, Yichao, et al.
Veröffentlicht: (2026)
von: Yuan, Yichao, et al.
Veröffentlicht: (2026)
Nalar: An agent serving framework
von: Laju, Marco, et al.
Veröffentlicht: (2026)
von: Laju, Marco, et al.
Veröffentlicht: (2026)
MoDM: Efficient Serving for Image Generation via Mixture-of-Diffusion Models
von: Xia, Yuchen, et al.
Veröffentlicht: (2025)
von: Xia, Yuchen, et al.
Veröffentlicht: (2025)
Understanding the Performance and Power of LLM Inferencing on Edge Accelerators
von: Arya, Mayank, et al.
Veröffentlicht: (2025)
von: Arya, Mayank, et al.
Veröffentlicht: (2025)
SCOOT: SLO-Oriented Performance Tuning for LLM Inference Engines
von: Cheng, Ke, et al.
Veröffentlicht: (2024)
von: Cheng, Ke, et al.
Veröffentlicht: (2024)
RIPPLE++: An Incremental Framework for Efficient GNN Inference on Evolving Graphs
von: Naman, Pranjal, et al.
Veröffentlicht: (2026)
von: Naman, Pranjal, et al.
Veröffentlicht: (2026)
CROWDio: A Practical Mobile Crowd Computing Framework with Developer-Oriented Design, Adaptive Scheduling, and Fault Resilience
von: Manamperi, Lakshani, et al.
Veröffentlicht: (2026)
von: Manamperi, Lakshani, et al.
Veröffentlicht: (2026)
Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads
von: Hu, Cunchen, et al.
Veröffentlicht: (2024)
von: Hu, Cunchen, et al.
Veröffentlicht: (2024)
The Merit of Simple Policies: Buying Performance With Parallelism and System Architecture
von: Yildiz, Mert, et al.
Veröffentlicht: (2025)
von: Yildiz, Mert, et al.
Veröffentlicht: (2025)
Byzantine Fault-Tolerant Min-Max Optimization
von: Liu, Shuo, et al.
Veröffentlicht: (2022)
von: Liu, Shuo, et al.
Veröffentlicht: (2022)
Assortment of Attention Heads: Accelerating Federated PEFT with Head Pruning and Strategic Client Selection
von: Venkatesha, Yeshwanth, et al.
Veröffentlicht: (2025)
von: Venkatesha, Yeshwanth, et al.
Veröffentlicht: (2025)
Kavier: Exploring Performance, Sustainability, and Efficiency of LLM Ecosystems under Inference through Cache-Aware Discrete-Event Simulation
von: Nicolae, Radu, et al.
Veröffentlicht: (2026)
von: Nicolae, Radu, et al.
Veröffentlicht: (2026)
Dispatching Odyssey: Exploring Performance in Computing Clusters under Real-world Workloads
von: Yildiz, Mert, et al.
Veröffentlicht: (2025)
von: Yildiz, Mert, et al.
Veröffentlicht: (2025)
RAPID-LLM: Resilience-Aware Performance analysis of Infrastructure for Distributed LLM Training and Inference
von: Karfakis, George, et al.
Veröffentlicht: (2025)
von: Karfakis, George, et al.
Veröffentlicht: (2025)
Cloud Native System for LLM Inference Serving
von: Xu, Minxian, et al.
Veröffentlicht: (2025)
von: Xu, Minxian, et al.
Veröffentlicht: (2025)
Enabling Dynamic Sparsity in Quantized LLM Inference
von: Wang, Rongxiang, et al.
Veröffentlicht: (2025)
von: Wang, Rongxiang, et al.
Veröffentlicht: (2025)
Federated Inference for Heterogeneous LLM Communication and Collaboration
von: Chen, Zihan, et al.
Veröffentlicht: (2026)
von: Chen, Zihan, et al.
Veröffentlicht: (2026)
Distributed Inference Performance Optimization for LLMs on CPUs
von: He, Pujiang, et al.
Veröffentlicht: (2024)
von: He, Pujiang, et al.
Veröffentlicht: (2024)
Fast and Cost-effective Speculative Edge-Cloud Decoding with Early Exits
von: Venkatesha, Yeshwanth, et al.
Veröffentlicht: (2025)
von: Venkatesha, Yeshwanth, et al.
Veröffentlicht: (2025)
From Attention to Disaggregation: Tracing the Evolution of LLM Inference
von: Kumar, Madabattula Rajesh, et al.
Veröffentlicht: (2025)
von: Kumar, Madabattula Rajesh, et al.
Veröffentlicht: (2025)
Towards Resource-Efficient Serverless LLM Inference with SLINFER
von: Xu, Chuhao, et al.
Veröffentlicht: (2025)
von: Xu, Chuhao, et al.
Veröffentlicht: (2025)
Argus: Token Aware Distributed LLM Inference Optimization
von: Wu, Panlong, et al.
Veröffentlicht: (2025)
von: Wu, Panlong, et al.
Veröffentlicht: (2025)
Toward Sustainability-Aware LLM Inference on Edge Clusters
von: Rajashekar, Kolichala, et al.
Veröffentlicht: (2025)
von: Rajashekar, Kolichala, et al.
Veröffentlicht: (2025)
Distributed On-Device LLM Inference With Over-the-Air Computation
von: Zhang, Kai, et al.
Veröffentlicht: (2025)
von: Zhang, Kai, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
On Evaluating Performance of LLM Inference Serving Systems
von: Agrawal, Amey, et al.
Veröffentlicht: (2025) -
Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve
von: Agrawal, Amey, et al.
Veröffentlicht: (2024) -
TokenWeave: Efficient Compute-Communication Overlap for Distributed LLM Inference
von: Gond, Raja, et al.
Veröffentlicht: (2025) -
Niyama : Breaking the Silos of LLM Inference Serving
von: Goel, Kanishk, et al.
Veröffentlicht: (2025) -
Kascade: A Practical Sparse Attention Method for Long-Context LLM Inference
von: Deshmukh, Dhruv, et al.
Veröffentlicht: (2025)