Meta-Metrics and Best Practices for System-Level Inference Performance Benchmarking
Fuente:
arXiv
Saved in:
| Main Authors: | Salaria, Shweta, Liu, Zhuoran, Gonzalez, Nelson Mimura |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Statistical Modeling and Uncertainty Estimation of LLM Inference Systems
by: Ray, Kaustabha, et al.
Published: (2025)
by: Ray, Kaustabha, et al.
Published: (2025)
Interpreting Performance Profiles with Deep Learning
by: Liu, Zhuoran
Published: (2025)
by: Liu, Zhuoran
Published: (2025)
Dissecting Embedding Bag Performance in DLRM Inference
by: Ambati, Chandrish, et al.
Published: (2025)
by: Ambati, Chandrish, et al.
Published: (2025)
Characterize LSM-tree Compaction Performance via On-Device LLM Inference
by: Ding, Jiabiao, et al.
Published: (2026)
by: Ding, Jiabiao, et al.
Published: (2026)
A Continuous Benchmarking Infrastructure for High-Performance Computing Applications
by: Alt, Christoph, et al.
Published: (2024)
by: Alt, Christoph, et al.
Published: (2024)
KG-EDAS: A Meta-Metric Framework for Evaluating Knowledge Graph Completion Models
by: Gul, Haji, et al.
Published: (2025)
by: Gul, Haji, et al.
Published: (2025)
Explainable Port Mapping Inference with Sparse Performance Counters for AMD's Zen Architectures
by: Ritter, Fabian, et al.
Published: (2024)
by: Ritter, Fabian, et al.
Published: (2024)
Towards CPU Performance Prediction: New Challenge Benchmark Dataset and Novel Approach
by: Liu, Xiaoman
Published: (2024)
by: Liu, Xiaoman
Published: (2024)
AI Application Benchmarking: Power-Aware Performance Analysis for Vision and Language Models
by: Mayr, Martin, et al.
Published: (2026)
by: Mayr, Martin, et al.
Published: (2026)
Unikernels vs. Containers: A Runtime-Level Performance Comparison for Resource-Constrained Edge Workloads
by: Dinh-Tuan, Hai
Published: (2025)
by: Dinh-Tuan, Hai
Published: (2025)
LLM Swiss Round: Aggregating Multi-Benchmark Performance via Competitive Swiss-System Dynamics
by: Liu, Jiashuo, et al.
Published: (2025)
by: Liu, Jiashuo, et al.
Published: (2025)
Dissecting RISC-V Performance: Practical PMU Profiling and Hardware-Agnostic Roofline Analysis on Emerging Platforms
by: Batashev, Alexander
Published: (2025)
by: Batashev, Alexander
Published: (2025)
Forecasting GPU Performance for Deep Learning Training and Inference
by: Lee, Seonho, et al.
Published: (2024)
by: Lee, Seonho, et al.
Published: (2024)
Pinching-Antenna Systems For Indoor Immersive Communications: A 3D-Modeling Based Performance Analysis
by: Wang, Yulei, et al.
Published: (2025)
by: Wang, Yulei, et al.
Published: (2025)
XRFlux: Virtual Reality Benchmark for Edge Caching Systems
by: Alfares, Nader, et al.
Published: (2024)
by: Alfares, Nader, et al.
Published: (2024)
Two Criteria for Performance Analysis of Optimization Algorithms
by: Jing, Yunpeng, et al.
Published: (2024)
by: Jing, Yunpeng, et al.
Published: (2024)
Towards Computational Performance Engineering for Unsupervised Concept Drift Detection -- Complexities, Benchmarking, Performance Analysis
by: Werner, Elias, et al.
Published: (2023)
by: Werner, Elias, et al.
Published: (2023)
Gem5-AcceSys: Enabling System-Level Exploration of Standard Interconnects for Novel Accelerators
by: Liu, Qunyou, et al.
Published: (2025)
by: Liu, Qunyou, et al.
Published: (2025)
Performance of Confidential Computing GPUs
by: Ibarra, Antonio Martínez, et al.
Published: (2025)
by: Ibarra, Antonio Martínez, et al.
Published: (2025)
Hierarchical Analyses Applied to Computer System Performance: Review and Call for Further Studies
by: Thomasian, Alexander
Published: (2024)
by: Thomasian, Alexander
Published: (2024)
COMPASS: A Unified Decision-Intelligence System for Navigating Performance Trade-off in HPC
by: Lahiry, Ankur, et al.
Published: (2026)
by: Lahiry, Ankur, et al.
Published: (2026)
An Empirical Study on Method-Level Performance Evolution in Open-Source Java Projects
by: Shahedi, Kaveh, et al.
Published: (2025)
by: Shahedi, Kaveh, et al.
Published: (2025)
Inference performance evaluation for LLMs on edge devices with a novel benchmarking framework and metric
by: Chen, Hao, et al.
Published: (2025)
by: Chen, Hao, et al.
Published: (2025)
Understanding Power Consumption Metric on Heterogeneous Memory Systems
by: Proaño, Andrès Rubio, et al.
Published: (2024)
by: Proaño, Andrès Rubio, et al.
Published: (2024)
MambaCPU: Enhanced Correlation Mining with State Space Models for CPU Performance Prediction
by: Liu, Xiaoman
Published: (2024)
by: Liu, Xiaoman
Published: (2024)
MoE-Inference-Bench: Performance Evaluation of Mixture of Expert Large Language and Vision Models
by: Chitty-Venkata, Krishna Teja, et al.
Published: (2025)
by: Chitty-Venkata, Krishna Teja, et al.
Published: (2025)
Columbo: Low Level End-to-End System Traces through Modular Full-System Simulation
by: Görgen, Jakob, et al.
Published: (2024)
by: Görgen, Jakob, et al.
Published: (2024)
Updates on the Low-Level Abstraction of Memory Access
by: Gruber, Bernhard Manfred
Published: (2023)
by: Gruber, Bernhard Manfred
Published: (2023)
PerfSeer: An Efficient and Accurate Deep Learning Models Performance Predictor
by: Zhao, Xinlong, et al.
Published: (2025)
by: Zhao, Xinlong, et al.
Published: (2025)
Accurate Performance Modeling And Uncertainty Analysis of Lossy Compression in Scientific Applications
by: Liu, Youyuan, et al.
Published: (2024)
by: Liu, Youyuan, et al.
Published: (2024)
Performance Characterization of Expert Router for Scalable LLM Inference
by: Pichlmeier, Josef, et al.
Published: (2024)
by: Pichlmeier, Josef, et al.
Published: (2024)
Adaptive Orchestration for Large-Scale Inference on Heterogeneous Accelerator Systems Balancing Cost, Performance, and Resilience
by: Biran, Yahav, et al.
Published: (2025)
by: Biran, Yahav, et al.
Published: (2025)
RAGPerf: An End-to-End Benchmarking Framework for Retrieval-Augmented Generation Systems
by: Li, Shaobo, et al.
Published: (2026)
by: Li, Shaobo, et al.
Published: (2026)
Machine Learning-Guided Memory Optimization for DLRM Inference on Tiered Memory
by: Ren, Jie, et al.
Published: (2025)
by: Ren, Jie, et al.
Published: (2025)
Profiling Large Language Model Inference on Apple Silicon: A Quantization Perspective
by: Benazir, Afsara, et al.
Published: (2025)
by: Benazir, Afsara, et al.
Published: (2025)
A System-Level Dynamic Binary Translator using Automatically-Learned Translation Rules
by: Jiang, Jinhu, et al.
Published: (2024)
by: Jiang, Jinhu, et al.
Published: (2024)
SparseInfer: Training-free Prediction of Activation Sparsity for Fast LLM Inference
by: Shin, Jiho, et al.
Published: (2024)
by: Shin, Jiho, et al.
Published: (2024)
ADS Performance Revisited
by: Weber, Alexander, et al.
Published: (2024)
by: Weber, Alexander, et al.
Published: (2024)
Wasure: A Modular Toolkit for Comprehensive WebAssembly Benchmarking
by: Carissimi, Riccardo, et al.
Published: (2026)
by: Carissimi, Riccardo, et al.
Published: (2026)
CAPSim: A Fast CPU Performance Simulator Using Attention-based Predictor
by: Xu, Buqing, et al.
Published: (2025)
by: Xu, Buqing, et al.
Published: (2025)
Similar Items
-
Statistical Modeling and Uncertainty Estimation of LLM Inference Systems
by: Ray, Kaustabha, et al.
Published: (2025) -
Interpreting Performance Profiles with Deep Learning
by: Liu, Zhuoran
Published: (2025) -
Dissecting Embedding Bag Performance in DLRM Inference
by: Ambati, Chandrish, et al.
Published: (2025) -
Characterize LSM-tree Compaction Performance via On-Device LLM Inference
by: Ding, Jiabiao, et al.
Published: (2026) -
A Continuous Benchmarking Infrastructure for High-Performance Computing Applications
by: Alt, Christoph, et al.
Published: (2024)