How Usable is Automated Feature Engineering for Tabular Data?

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Schäfer, Bastian, Purucker, Lennart, Janowski, Maciej, Hutter, Frank
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912544459849728
author Schäfer, Bastian
Purucker, Lennart
Janowski, Maciej
Hutter, Frank
author_facet Schäfer, Bastian
Purucker, Lennart
Janowski, Maciej
Hutter, Frank
contents Tabular data, consisting of rows and columns, is omnipresent across various machine learning applications. Each column represents a feature, and features can be combined or transformed to create new, more informative features. Such feature engineering is essential to achieve peak performance in machine learning. Since manual feature engineering is expensive and time-consuming, a substantial effort has been put into automating it. Yet, existing automated feature engineering (AutoFE) methods have never been investigated regarding their usability for practitioners. Thus, we investigated 53 AutoFE methods. We found that these methods are, in general, hard to use, lack documentation, and have no active communities. Furthermore, no method allows users to set time and memory constraints, which we see as a necessity for usable automation. Our survey highlights the need for future work on usable, well-engineered AutoFE methods.
format Preprint
id arxiv_https___arxiv_org_abs_2508_13932
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle How Usable is Automated Feature Engineering for Tabular Data?
Schäfer, Bastian
Purucker, Lennart
Janowski, Maciej
Hutter, Frank
Machine Learning
Tabular data, consisting of rows and columns, is omnipresent across various machine learning applications. Each column represents a feature, and features can be combined or transformed to create new, more informative features. Such feature engineering is essential to achieve peak performance in machine learning. Since manual feature engineering is expensive and time-consuming, a substantial effort has been put into automating it. Yet, existing automated feature engineering (AutoFE) methods have never been investigated regarding their usability for practitioners. Thus, we investigated 53 AutoFE methods. We found that these methods are, in general, hard to use, lack documentation, and have no active communities. Furthermore, no method allows users to set time and memory constraints, which we see as a necessity for usable automation. Our survey highlights the need for future work on usable, well-engineered AutoFE methods.
title How Usable is Automated Feature Engineering for Tabular Data?
topic Machine Learning
url https://arxiv.org/abs/2508.13932