Calibrated Dataset Condensation for Faster Hyperparameter Search

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ding, Mucong, Xu, Yuancheng, Rabbani, Tahseen, Liu, Xiaoyu, Gravelle, Brian, Ranadive, Teresa, Tuan, Tai-Ching, Huang, Furong
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910460534587392
author Ding, Mucong
Xu, Yuancheng
Rabbani, Tahseen
Liu, Xiaoyu
Gravelle, Brian
Ranadive, Teresa
Tuan, Tai-Ching
Huang, Furong
author_facet Ding, Mucong
Xu, Yuancheng
Rabbani, Tahseen
Liu, Xiaoyu
Gravelle, Brian
Ranadive, Teresa
Tuan, Tai-Ching
Huang, Furong
contents Dataset condensation can be used to reduce the computational cost of training multiple models on a large dataset by condensing the training dataset into a small synthetic set. State-of-the-art approaches rely on matching the model gradients between the real and synthetic data. However, there is no theoretical guarantee of the generalizability of the condensed data: data condensation often generalizes poorly across hyperparameters/architectures in practice. This paper considers a different condensation objective specifically geared toward hyperparameter search. We aim to generate a synthetic validation dataset so that the validation-performance rankings of the models, with different hyperparameters, on the condensed and original datasets are comparable. We propose a novel hyperparameter-calibrated dataset condensation (HCDC) algorithm, which obtains the synthetic validation dataset by matching the hyperparameter gradients computed via implicit differentiation and efficient inverse Hessian approximation. Experiments demonstrate that the proposed framework effectively maintains the validation-performance rankings of models and speeds up hyperparameter/architecture search for tasks on both images and graphs.
format Preprint
id arxiv_https___arxiv_org_abs_2405_17535
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Calibrated Dataset Condensation for Faster Hyperparameter Search
Ding, Mucong
Xu, Yuancheng
Rabbani, Tahseen
Liu, Xiaoyu
Gravelle, Brian
Ranadive, Teresa
Tuan, Tai-Ching
Huang, Furong
Machine Learning
Artificial Intelligence
Dataset condensation can be used to reduce the computational cost of training multiple models on a large dataset by condensing the training dataset into a small synthetic set. State-of-the-art approaches rely on matching the model gradients between the real and synthetic data. However, there is no theoretical guarantee of the generalizability of the condensed data: data condensation often generalizes poorly across hyperparameters/architectures in practice. This paper considers a different condensation objective specifically geared toward hyperparameter search. We aim to generate a synthetic validation dataset so that the validation-performance rankings of the models, with different hyperparameters, on the condensed and original datasets are comparable. We propose a novel hyperparameter-calibrated dataset condensation (HCDC) algorithm, which obtains the synthetic validation dataset by matching the hyperparameter gradients computed via implicit differentiation and efficient inverse Hessian approximation. Experiments demonstrate that the proposed framework effectively maintains the validation-performance rankings of models and speeds up hyperparameter/architecture search for tasks on both images and graphs.
title Calibrated Dataset Condensation for Faster Hyperparameter Search
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2405.17535