Plug-and-Play Performance Estimation for LLM Services without Relying on Labeled Data
Fuente:
arXiv
Guardado en:
| Autores principales: | Wang, Can, Sui, Dianbo, Sun, Hongliang, Ding, Hao, Zhang, Bolin, Tu, Zhiying |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
CEBench: A Benchmarking Toolkit for the Cost-Effectiveness of LLM Pipelines
por: Sun, Wenbo, et al.
Publicado: (2024)
por: Sun, Wenbo, et al.
Publicado: (2024)
KernelBenchX: A Comprehensive Benchmark for Evaluating LLM-Generated GPU Kernels
por: Wang, Han, et al.
Publicado: (2026)
por: Wang, Han, et al.
Publicado: (2026)
VDTuner: Automated Performance Tuning for Vector Data Management Systems
por: Yang, Tiannuo, et al.
Publicado: (2024)
por: Yang, Tiannuo, et al.
Publicado: (2024)
lm-Meter: Unveiling Runtime Inference Latency for On-Device Language Models
por: Wang, Haoxin, et al.
Publicado: (2025)
por: Wang, Haoxin, et al.
Publicado: (2025)
MarginGate: Sparse Margin-Triggered Verification for Batch-Invariant LLM Inference
por: Chu, Kexin, et al.
Publicado: (2026)
por: Chu, Kexin, et al.
Publicado: (2026)
Ragged Paged Attention: A High-Performance and Flexible LLM Inference Kernel for TPU
por: Jiang, Jevin, et al.
Publicado: (2026)
por: Jiang, Jevin, et al.
Publicado: (2026)
Application Research On Real-Time Perception Of Device Performance Status
por: Wang, Zhe, et al.
Publicado: (2024)
por: Wang, Zhe, et al.
Publicado: (2024)
Efficient Graph Knowledge Distillation from GNNs to Kolmogorov--Arnold Networks via Self-Attention Dynamic Sampling
por: Cui, Can, et al.
Publicado: (2025)
por: Cui, Can, et al.
Publicado: (2025)
LLM Swiss Round: Aggregating Multi-Benchmark Performance via Competitive Swiss-System Dynamics
por: Liu, Jiashuo, et al.
Publicado: (2025)
por: Liu, Jiashuo, et al.
Publicado: (2025)
Towards Computational Performance Engineering for Unsupervised Concept Drift Detection -- Complexities, Benchmarking, Performance Analysis
por: Werner, Elias, et al.
Publicado: (2023)
por: Werner, Elias, et al.
Publicado: (2023)
oneDNN Graph Compiler: A Hybrid Approach for High-Performance Deep Learning Compilation
por: Li, Jianhui, et al.
Publicado: (2023)
por: Li, Jianhui, et al.
Publicado: (2023)
An Interpretable Latency Model for Speculative Decoding in LLM Serving
por: Kong, Linghao, et al.
Publicado: (2026)
por: Kong, Linghao, et al.
Publicado: (2026)
An Inquiry into Datacenter TCO for LLM Inference with FP8
por: Kim, Jiwoo, et al.
Publicado: (2025)
por: Kim, Jiwoo, et al.
Publicado: (2025)
Forecasting GPU Performance for Deep Learning Training and Inference
por: Lee, Seonho, et al.
Publicado: (2024)
por: Lee, Seonho, et al.
Publicado: (2024)
PARD: Accelerating LLM Inference with Low-Cost PARallel Draft Model Adaptation
por: An, Zihao, et al.
Publicado: (2025)
por: An, Zihao, et al.
Publicado: (2025)
Performance Modeling of Data Storage Systems using Generative Models
por: Al-Maeeni, Abdalaziz Rashid, et al.
Publicado: (2023)
por: Al-Maeeni, Abdalaziz Rashid, et al.
Publicado: (2023)
Towards an Integrated Performance Framework for Fire Science and Management Workflows
por: Ahmed, H., et al.
Publicado: (2024)
por: Ahmed, H., et al.
Publicado: (2024)
CPINN-ABPI: Physics-Informed Neural Networks for Accurate Power Estimation in MPSoCs
por: Elshamy, Mohamed R., et al.
Publicado: (2025)
por: Elshamy, Mohamed R., et al.
Publicado: (2025)
A Framework for Effective Invocation Methods of Various LLM Services
por: Wang, Can, et al.
Publicado: (2024)
por: Wang, Can, et al.
Publicado: (2024)
CirrusBench: Evaluating LLM-based Agents Beyond Correctness in Real-World Cloud Service Environments
por: Yu, Yi, et al.
Publicado: (2026)
por: Yu, Yi, et al.
Publicado: (2026)
Ensuring Reliability of Curated EHR-Derived Data: The Validation of Accuracy for LLM/ML-Extracted Information and Data (VALID) Framework
por: Estevez, Melissa, et al.
Publicado: (2025)
por: Estevez, Melissa, et al.
Publicado: (2025)
GreenServ: Energy-Efficient Context-Aware Dynamic Routing for Multi-Model LLM Inference
por: Ziller, Thomas, et al.
Publicado: (2026)
por: Ziller, Thomas, et al.
Publicado: (2026)
V-Seek: Accelerating LLM Reasoning on Open-hardware Server-class RISC-V Platforms
por: Rodrigo, Javier J. Poveda, et al.
Publicado: (2025)
por: Rodrigo, Javier J. Poveda, et al.
Publicado: (2025)
A Kernel-Based Approach for Accurate Steady-State Detection in Performance Time Series
por: Beseda, Martin, et al.
Publicado: (2025)
por: Beseda, Martin, et al.
Publicado: (2025)
MoE-Inference-Bench: Performance Evaluation of Mixture of Expert Large Language and Vision Models
por: Chitty-Venkata, Krishna Teja, et al.
Publicado: (2025)
por: Chitty-Venkata, Krishna Teja, et al.
Publicado: (2025)
SAfEPaTh: A System-Level Approach for Efficient Power and Thermal Estimation of Convolutional Neural Network Accelerator
por: Chen, Yukai, et al.
Publicado: (2024)
por: Chen, Yukai, et al.
Publicado: (2024)
Reducing Latency of LLM Search Agent via Speculation-based Algorithm-System Co-Design
por: Huang, Zixiao, et al.
Publicado: (2025)
por: Huang, Zixiao, et al.
Publicado: (2025)
Single-Thread JPEG Decoder Benchmarks Mis-Evaluate ML Data Loaders
por: Iglovikov, Vladimir, et al.
Publicado: (2026)
por: Iglovikov, Vladimir, et al.
Publicado: (2026)
Large-Scale Data Parallelization of Product Quantization and Inverted Indexing Using Dask
por: Abraham, Ashley N., et al.
Publicado: (2026)
por: Abraham, Ashley N., et al.
Publicado: (2026)
Phantora: Maximizing Code Reuse in Simulation-based Machine Learning System Performance Estimation
por: Qin, Jianxing, et al.
Publicado: (2025)
por: Qin, Jianxing, et al.
Publicado: (2025)
Development and Comparative Evaluation of Three Artificial Intelligence Models (NLP, LLM, JEPA) for Predicting Triage in Emergency Departments: A 7-Month Retrospective Proof-of-Concept
por: Lansiaux, Edouard, et al.
Publicado: (2025)
por: Lansiaux, Edouard, et al.
Publicado: (2025)
SynthEval: A Framework for Detailed Utility and Privacy Evaluation of Tabular Synthetic Data
por: Lautrup, Anton Danholt, et al.
Publicado: (2024)
por: Lautrup, Anton Danholt, et al.
Publicado: (2024)
APOLLO: SGD-like Memory, AdamW-level Performance
por: Zhu, Hanqing, et al.
Publicado: (2024)
por: Zhu, Hanqing, et al.
Publicado: (2024)
EDGC: Entropy-driven Dynamic Gradient Compression for Efficient LLM Training
por: Yi, Qingao, et al.
Publicado: (2025)
por: Yi, Qingao, et al.
Publicado: (2025)
ClusterKV: Manipulating LLM KV Cache in Semantic Space for Recallable Compression
por: Liu, Guangda, et al.
Publicado: (2024)
por: Liu, Guangda, et al.
Publicado: (2024)
Systematic Characterization of LLM Quantization: A Performance, Energy, and Quality Perspective
por: Shi, Tianyao, et al.
Publicado: (2025)
por: Shi, Tianyao, et al.
Publicado: (2025)
DistZO2: High-Throughput and Memory-Efficient Zeroth-Order Fine-tuning LLMs with Distributed Parallel Computing
por: Wang, Liangyu, et al.
Publicado: (2025)
por: Wang, Liangyu, et al.
Publicado: (2025)
Forecasting LLM Inference Performance via Hardware-Agnostic Analytical Modeling
por: Patwari, Rajeev, et al.
Publicado: (2025)
por: Patwari, Rajeev, et al.
Publicado: (2025)
Confidential LLM Inference: Performance and Cost Across CPU and GPU TEEs
por: Chrapek, Marcin, et al.
Publicado: (2025)
por: Chrapek, Marcin, et al.
Publicado: (2025)
Neuralink: Fast LLM Inference on Smartphones with Neuron Co-Activation Linking
por: Wang, Tuowei, et al.
Publicado: (2024)
por: Wang, Tuowei, et al.
Publicado: (2024)
Ejemplares similares
-
CEBench: A Benchmarking Toolkit for the Cost-Effectiveness of LLM Pipelines
por: Sun, Wenbo, et al.
Publicado: (2024) -
KernelBenchX: A Comprehensive Benchmark for Evaluating LLM-Generated GPU Kernels
por: Wang, Han, et al.
Publicado: (2026) -
VDTuner: Automated Performance Tuning for Vector Data Management Systems
por: Yang, Tiannuo, et al.
Publicado: (2024) -
lm-Meter: Unveiling Runtime Inference Latency for On-Device Language Models
por: Wang, Haoxin, et al.
Publicado: (2025) -
MarginGate: Sparse Margin-Triggered Verification for Batch-Invariant LLM Inference
por: Chu, Kexin, et al.
Publicado: (2026)