TabDPT: Scaling Tabular Foundation Models on Real Data

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Ma, Junwei, Thomas, Valentin, Hosseinzadeh, Rasa, Labach, Alex, Kamkari, Hamidreza, Cresswell, Jesse C., Golestan, Keyvan, Yu, Guangwei, Caterini, Anthony L., Volkovs, Maksims
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908771304865792
author Ma, Junwei
Thomas, Valentin
Hosseinzadeh, Rasa
Labach, Alex
Kamkari, Hamidreza
Cresswell, Jesse C.
Golestan, Keyvan
Yu, Guangwei
Caterini, Anthony L.
Volkovs, Maksims
author_facet Ma, Junwei
Thomas, Valentin
Hosseinzadeh, Rasa
Labach, Alex
Kamkari, Hamidreza
Cresswell, Jesse C.
Golestan, Keyvan
Yu, Guangwei
Caterini, Anthony L.
Volkovs, Maksims
contents Tabular data is one of the most ubiquitous sources of information worldwide, spanning a wide variety of domains. This inherent heterogeneity has slowed the development of Tabular Foundation Models (TFMs) capable of fast generalization to unseen datasets. In-Context Learning (ICL) has recently emerged as a promising solution for TFMs, enabling dynamic adaptation to new tasks without additional tuning. While many studies have attempted to re-purpose large language models for tabular ICL, they have had limited success, so recent works have focused on developing tabular-specific foundation models. In this work, we propose an approach to combine ICL-based retrieval with self supervised learning to train tabular foundation models. We also investigate the utility of real vs. synthetic data for model pre-training, and show that real data can contain useful signal not easily captured in synthetic training. Specifically, we show that incorporating real data during the pre-training phase can lead to significantly faster training and better downstream generalization to unseen data. Our resulting model, TabDPT, achieves strong performance on both regression (CTR23) and classification (CC18) benchmarks. Importantly, we also demonstrate that with our pre-training procedure, scaling both model and data size leads to consistent performance improvements that follow power laws. This echoes scaling laws in LLMs and other foundation models, and suggests that large-scale TFMs can be achievable. We open-source our full pipeline: inference code including trained model weights can be found at github.com/layer6ai-labs/TabDPT-inference, and the training code to reproduce experiments can be found at github.com/layer6ai-labs/TabDPT-training.
format Preprint
id arxiv_https___arxiv_org_abs_2410_18164
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle TabDPT: Scaling Tabular Foundation Models on Real Data
Ma, Junwei
Thomas, Valentin
Hosseinzadeh, Rasa
Labach, Alex
Kamkari, Hamidreza
Cresswell, Jesse C.
Golestan, Keyvan
Yu, Guangwei
Caterini, Anthony L.
Volkovs, Maksims
Machine Learning
Artificial Intelligence
Tabular data is one of the most ubiquitous sources of information worldwide, spanning a wide variety of domains. This inherent heterogeneity has slowed the development of Tabular Foundation Models (TFMs) capable of fast generalization to unseen datasets. In-Context Learning (ICL) has recently emerged as a promising solution for TFMs, enabling dynamic adaptation to new tasks without additional tuning. While many studies have attempted to re-purpose large language models for tabular ICL, they have had limited success, so recent works have focused on developing tabular-specific foundation models. In this work, we propose an approach to combine ICL-based retrieval with self supervised learning to train tabular foundation models. We also investigate the utility of real vs. synthetic data for model pre-training, and show that real data can contain useful signal not easily captured in synthetic training. Specifically, we show that incorporating real data during the pre-training phase can lead to significantly faster training and better downstream generalization to unseen data. Our resulting model, TabDPT, achieves strong performance on both regression (CTR23) and classification (CC18) benchmarks. Importantly, we also demonstrate that with our pre-training procedure, scaling both model and data size leads to consistent performance improvements that follow power laws. This echoes scaling laws in LLMs and other foundation models, and suggests that large-scale TFMs can be achievable. We open-source our full pipeline: inference code including trained model weights can be found at github.com/layer6ai-labs/TabDPT-inference, and the training code to reproduce experiments can be found at github.com/layer6ai-labs/TabDPT-training.
title TabDPT: Scaling Tabular Foundation Models on Real Data
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2410.18164