Frozen Layers: Memory-efficient Many-fidelity Hyperparameter Optimization

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Carstensen, Timur, Mallik, Neeratyoy, Hutter, Frank, Rapp, Martin
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866910913370521600
author Carstensen, Timur
Mallik, Neeratyoy
Hutter, Frank
Rapp, Martin
author_facet Carstensen, Timur
Mallik, Neeratyoy
Hutter, Frank
Rapp, Martin
contents As model sizes grow, finding efficient and cost-effective hyperparameter optimization (HPO) methods becomes increasingly crucial for deep learning pipelines. While multi-fidelity HPO (MF-HPO) trades off computational resources required for DL training with lower fidelity estimations, existing fidelity sources often fail under lower compute and memory constraints. We propose a novel fidelity source: the number of layers that are trained or frozen during training. For deep networks, this approach offers significant compute and memory savings while preserving rank correlations between hyperparameters at low fidelities compared to full model training. We demonstrate this in our empirical evaluation across ResNets and Transformers and additionally analyze the utility of frozen layers as a fidelity in using GPU resources as a fidelity in HPO, and for a combined MF-HPO with other fidelity sources. This contribution opens new applications for MF-HPO with hardware resources as a fidelity and creates opportunities for improved algorithms navigating joint fidelity spaces.
format Preprint
id arxiv_https___arxiv_org_abs_2504_10735
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Frozen Layers: Memory-efficient Many-fidelity Hyperparameter Optimization
Carstensen, Timur
Mallik, Neeratyoy
Hutter, Frank
Rapp, Martin
Machine Learning
Artificial Intelligence
As model sizes grow, finding efficient and cost-effective hyperparameter optimization (HPO) methods becomes increasingly crucial for deep learning pipelines. While multi-fidelity HPO (MF-HPO) trades off computational resources required for DL training with lower fidelity estimations, existing fidelity sources often fail under lower compute and memory constraints. We propose a novel fidelity source: the number of layers that are trained or frozen during training. For deep networks, this approach offers significant compute and memory savings while preserving rank correlations between hyperparameters at low fidelities compared to full model training. We demonstrate this in our empirical evaluation across ResNets and Transformers and additionally analyze the utility of frozen layers as a fidelity in using GPU resources as a fidelity in HPO, and for a combined MF-HPO with other fidelity sources. This contribution opens new applications for MF-HPO with hardware resources as a fidelity and creates opportunities for improved algorithms navigating joint fidelity spaces.
title Frozen Layers: Memory-efficient Many-fidelity Hyperparameter Optimization
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2504.10735