Enhancing Tabular Data Optimization with a Flexible Graph-based Reinforced Exploration Strategy

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Xiaohan, Wang, Dongjie, Ning, Zhiyuan, Qiao, Ziyue, Long, Qingqing, Zhu, Haowei, Wu, Min, Zhou, Yuanchun, Xiao, Meng
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913386697064448
author Huang, Xiaohan
Wang, Dongjie
Ning, Zhiyuan
Qiao, Ziyue
Long, Qingqing
Zhu, Haowei
Wu, Min
Zhou, Yuanchun
Xiao, Meng
author_facet Huang, Xiaohan
Wang, Dongjie
Ning, Zhiyuan
Qiao, Ziyue
Long, Qingqing
Zhu, Haowei
Wu, Min
Zhou, Yuanchun
Xiao, Meng
contents Tabular data optimization methods aim to automatically find an optimal feature transformation process that generates high-value features and improves the performance of downstream machine learning tasks. Current frameworks for automated feature transformation rely on iterative sequence generation tasks, optimizing decision strategies through performance feedback from downstream tasks. However, these approaches fail to effectively utilize historical decision-making experiences and overlook potential relationships among generated features, thus limiting the depth of knowledge extraction. Moreover, the granularity of the decision-making process lacks dynamic backtracking capabilities for individual features, leading to insufficient adaptability when encountering inefficient pathways, adversely affecting overall robustness and exploration efficiency. To address the limitations observed in current automatic feature engineering frameworks, we introduce a novel method that utilizes a feature-state transformation graph to effectively preserve the entire feature transformation journey, where each node represents a specific transformation state. During exploration, three cascading agents iteratively select nodes and idea mathematical operations to generate new transformation states. This strategy leverages the inherent properties of the graph structure, allowing for the preservation and reuse of valuable transformations. It also enables backtracking capabilities through graph pruning techniques, which can rectify inefficient transformation paths. To validate the efficacy and flexibility of our approach, we conducted comprehensive experiments and detailed case studies, demonstrating superior performance in diverse scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2406_07404
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Enhancing Tabular Data Optimization with a Flexible Graph-based Reinforced Exploration Strategy
Huang, Xiaohan
Wang, Dongjie
Ning, Zhiyuan
Qiao, Ziyue
Long, Qingqing
Zhu, Haowei
Wu, Min
Zhou, Yuanchun
Xiao, Meng
Machine Learning
Tabular data optimization methods aim to automatically find an optimal feature transformation process that generates high-value features and improves the performance of downstream machine learning tasks. Current frameworks for automated feature transformation rely on iterative sequence generation tasks, optimizing decision strategies through performance feedback from downstream tasks. However, these approaches fail to effectively utilize historical decision-making experiences and overlook potential relationships among generated features, thus limiting the depth of knowledge extraction. Moreover, the granularity of the decision-making process lacks dynamic backtracking capabilities for individual features, leading to insufficient adaptability when encountering inefficient pathways, adversely affecting overall robustness and exploration efficiency. To address the limitations observed in current automatic feature engineering frameworks, we introduce a novel method that utilizes a feature-state transformation graph to effectively preserve the entire feature transformation journey, where each node represents a specific transformation state. During exploration, three cascading agents iteratively select nodes and idea mathematical operations to generate new transformation states. This strategy leverages the inherent properties of the graph structure, allowing for the preservation and reuse of valuable transformations. It also enables backtracking capabilities through graph pruning techniques, which can rectify inefficient transformation paths. To validate the efficacy and flexibility of our approach, we conducted comprehensive experiments and detailed case studies, demonstrating superior performance in diverse scenarios.
title Enhancing Tabular Data Optimization with a Flexible Graph-based Reinforced Exploration Strategy
topic Machine Learning
url https://arxiv.org/abs/2406.07404