Using Large Language Models for Hyperparameter Optimization

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zhang, Michael R., Desai, Nishkrit, Bae, Juhan, Lorraine, Jonathan, Ba, Jimmy
Format: Preprint
Veröffentlicht: 2023
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912113784520704
author Zhang, Michael R.
Desai, Nishkrit
Bae, Juhan
Lorraine, Jonathan
Ba, Jimmy
author_facet Zhang, Michael R.
Desai, Nishkrit
Bae, Juhan
Lorraine, Jonathan
Ba, Jimmy
contents This paper explores the use of foundational large language models (LLMs) in hyperparameter optimization (HPO). Hyperparameters are critical in determining the effectiveness of machine learning models, yet their optimization often relies on manual approaches in limited-budget settings. By prompting LLMs with dataset and model descriptions, we develop a methodology where LLMs suggest hyperparameter configurations, which are iteratively refined based on model performance. Our empirical evaluations on standard benchmarks reveal that within constrained search budgets, LLMs can match or outperform traditional HPO methods like Bayesian optimization across different models on standard benchmarks. Furthermore, we propose to treat the code specifying our model as a hyperparameter, which the LLM outputs and affords greater flexibility than existing HPO approaches.
format Preprint
id arxiv_https___arxiv_org_abs_2312_04528
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Using Large Language Models for Hyperparameter Optimization
Zhang, Michael R.
Desai, Nishkrit
Bae, Juhan
Lorraine, Jonathan
Ba, Jimmy
Machine Learning
Artificial Intelligence
This paper explores the use of foundational large language models (LLMs) in hyperparameter optimization (HPO). Hyperparameters are critical in determining the effectiveness of machine learning models, yet their optimization often relies on manual approaches in limited-budget settings. By prompting LLMs with dataset and model descriptions, we develop a methodology where LLMs suggest hyperparameter configurations, which are iteratively refined based on model performance. Our empirical evaluations on standard benchmarks reveal that within constrained search budgets, LLMs can match or outperform traditional HPO methods like Bayesian optimization across different models on standard benchmarks. Furthermore, we propose to treat the code specifying our model as a hyperparameter, which the LLM outputs and affords greater flexibility than existing HPO approaches.
title Using Large Language Models for Hyperparameter Optimization
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2312.04528