Bench360: Benchmarking Local LLM Inference from 360 Degrees

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Stuhlmann, Linus, Argerich, Mauricio Fadel, Fürst, Jonathan
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915727493038080
author Stuhlmann, Linus
Argerich, Mauricio Fadel
Fürst, Jonathan
author_facet Stuhlmann, Linus
Argerich, Mauricio Fadel
Fürst, Jonathan
contents Running LLMs locally has become increasingly common, but users face a complex design space across models, quantization levels, inference engines, and serving scenarios. Existing inference benchmarks are fragmented and focus on isolated goals, offering little guidance for practical deployments. We present Bench360, a framework for evaluating local LLM inference across tasks, usage patterns, and system metrics in one place. Bench360 supports custom tasks, integrates multiple inference engines and quantization formats, and reports both task quality and system behavior (latency, throughput, energy, startup time). We demonstrate it on four NLP tasks across three GPUs and four engines, showing how design choices shape efficiency and output quality. Results confirm that tradeoffs are substantial and configuration choices depend on specific workloads and constraints. There is no universal best option, underscoring the need for comprehensive, deployment-oriented benchmarks.
format Preprint
id arxiv_https___arxiv_org_abs_2511_16682
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Bench360: Benchmarking Local LLM Inference from 360 Degrees
Stuhlmann, Linus
Argerich, Mauricio Fadel
Fürst, Jonathan
Computation and Language
Artificial Intelligence
Machine Learning
Performance
Running LLMs locally has become increasingly common, but users face a complex design space across models, quantization levels, inference engines, and serving scenarios. Existing inference benchmarks are fragmented and focus on isolated goals, offering little guidance for practical deployments. We present Bench360, a framework for evaluating local LLM inference across tasks, usage patterns, and system metrics in one place. Bench360 supports custom tasks, integrates multiple inference engines and quantization formats, and reports both task quality and system behavior (latency, throughput, energy, startup time). We demonstrate it on four NLP tasks across three GPUs and four engines, showing how design choices shape efficiency and output quality. Results confirm that tradeoffs are substantial and configuration choices depend on specific workloads and constraints. There is no universal best option, underscoring the need for comprehensive, deployment-oriented benchmarks.
title Bench360: Benchmarking Local LLM Inference from 360 Degrees
topic Computation and Language
Artificial Intelligence
Machine Learning
Performance
url https://arxiv.org/abs/2511.16682