LLMSYS-HPOBench: Hyperparameter Optimization Benchmark Suite for Real-World LLM Systems

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Siyu, Ye, Yulong, Xiang, Zezhen, Chen, Pengzhou, Xiong, Gangda, Chen, Tao
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915995391623168
author Wu, Siyu
Ye, Yulong
Xiang, Zezhen
Chen, Pengzhou
Xiong, Gangda
Chen, Tao
author_facet Wu, Siyu
Ye, Yulong
Xiang, Zezhen
Chen, Pengzhou
Xiong, Gangda
Chen, Tao
contents Large Language Model (LLM) systems have been the frontier of AI in many application domains, leading to new challenges and opportunities for hyperparameter optimization (HPO) for the AutoML community. However, this type of system exhibits an unprecedented compound space of hyperparameter configuration from both the AI and non-AI components; rich and nonlinear implications from the fidelity factors; and diverse costs of measuring hyperparameter configurations, none of which have been fully captured in existing benchmarks. This paper presents the first (live) benchmark suite and datasets for HPO of real-world LLM systems, dubbed LLMSYS-HPOBench, covering data related to the inference objective values of hyperparameter configurations profiled from running the LLM systems. Currently, LLMSYS-HPOBench contains 364,450 hyperparameter configurations with a dimensionality of 12-23, 3-5 dimensions of fidelity factor leading to 932 settings, 3-9 inference objective metrics, and 2-10 cost metrics, together with generated logs from measuring the LLM systems. What we seek to advocate is not only a revalidation of the existing HPO algorithms over the frontier LLM systems, but also to provide an evolving platform for the AutoML community to explore new directions of research in this regard. The benchmark suite has been made available at: https://github.com/ideas-labo/llmsys-hpobench
format Preprint
id arxiv_https___arxiv_org_abs_2605_08305
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle LLMSYS-HPOBench: Hyperparameter Optimization Benchmark Suite for Real-World LLM Systems
Wu, Siyu
Ye, Yulong
Xiang, Zezhen
Chen, Pengzhou
Xiong, Gangda
Chen, Tao
Machine Learning
Artificial Intelligence
Computation and Language
Performance
Software Engineering
Large Language Model (LLM) systems have been the frontier of AI in many application domains, leading to new challenges and opportunities for hyperparameter optimization (HPO) for the AutoML community. However, this type of system exhibits an unprecedented compound space of hyperparameter configuration from both the AI and non-AI components; rich and nonlinear implications from the fidelity factors; and diverse costs of measuring hyperparameter configurations, none of which have been fully captured in existing benchmarks. This paper presents the first (live) benchmark suite and datasets for HPO of real-world LLM systems, dubbed LLMSYS-HPOBench, covering data related to the inference objective values of hyperparameter configurations profiled from running the LLM systems. Currently, LLMSYS-HPOBench contains 364,450 hyperparameter configurations with a dimensionality of 12-23, 3-5 dimensions of fidelity factor leading to 932 settings, 3-9 inference objective metrics, and 2-10 cost metrics, together with generated logs from measuring the LLM systems. What we seek to advocate is not only a revalidation of the existing HPO algorithms over the frontier LLM systems, but also to provide an evolving platform for the AutoML community to explore new directions of research in this regard. The benchmark suite has been made available at: https://github.com/ideas-labo/llmsys-hpobench
title LLMSYS-HPOBench: Hyperparameter Optimization Benchmark Suite for Real-World LLM Systems
topic Machine Learning
Artificial Intelligence
Computation and Language
Performance
Software Engineering
url https://arxiv.org/abs/2605.08305