ILAEDA: An Imitation Learning Based Approach for Automatic Exploratory Data Analysis

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Manatkar, Abhijit, Patel, Devarsh, Patel, Hima, Manwani, Naresh
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910651073429504
author Manatkar, Abhijit
Patel, Devarsh
Patel, Hima
Manwani, Naresh
author_facet Manatkar, Abhijit
Patel, Devarsh
Patel, Hima
Manwani, Naresh
contents Automating end-to-end Exploratory Data Analysis (AutoEDA) is a challenging open problem, often tackled through Reinforcement Learning (RL) by learning to predict a sequence of analysis operations (FILTER, GROUP, etc). Defining rewards for each operation is a challenging task and existing methods rely on various \emph{interestingness measures} to craft reward functions to capture the importance of each operation. In this work, we argue that not all of the essential features of what makes an operation important can be accurately captured mathematically using rewards. We propose an AutoEDA model trained through imitation learning from expert EDA sessions, bypassing the need for manually defined interestingness measures. Our method, based on generative adversarial imitation learning (GAIL), generalizes well across datasets, even with limited expert data. We also introduce a novel approach for generating synthetic EDA demonstrations for training. Our method outperforms the existing state-of-the-art end-to-end EDA approach on benchmarks by upto 3x, showing strong performance and generalization, while naturally capturing diverse interestingness measures in generated EDA sessions.
format Preprint
id arxiv_https___arxiv_org_abs_2410_11276
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle ILAEDA: An Imitation Learning Based Approach for Automatic Exploratory Data Analysis
Manatkar, Abhijit
Patel, Devarsh
Patel, Hima
Manwani, Naresh
Machine Learning
Artificial Intelligence
Databases
Automating end-to-end Exploratory Data Analysis (AutoEDA) is a challenging open problem, often tackled through Reinforcement Learning (RL) by learning to predict a sequence of analysis operations (FILTER, GROUP, etc). Defining rewards for each operation is a challenging task and existing methods rely on various \emph{interestingness measures} to craft reward functions to capture the importance of each operation. In this work, we argue that not all of the essential features of what makes an operation important can be accurately captured mathematically using rewards. We propose an AutoEDA model trained through imitation learning from expert EDA sessions, bypassing the need for manually defined interestingness measures. Our method, based on generative adversarial imitation learning (GAIL), generalizes well across datasets, even with limited expert data. We also introduce a novel approach for generating synthetic EDA demonstrations for training. Our method outperforms the existing state-of-the-art end-to-end EDA approach on benchmarks by upto 3x, showing strong performance and generalization, while naturally capturing diverse interestingness measures in generated EDA sessions.
title ILAEDA: An Imitation Learning Based Approach for Automatic Exploratory Data Analysis
topic Machine Learning
Artificial Intelligence
Databases
url https://arxiv.org/abs/2410.11276