Real-TabPFN: Improving Tabular Foundation Models via Continued Pre-training With Real-World Data

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Garg, Anurag, Ali, Muhammad, Hollmann, Noah, Purucker, Lennart, Müller, Samuel, Hutter, Frank
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909676012044288
author Garg, Anurag
Ali, Muhammad
Hollmann, Noah
Purucker, Lennart
Müller, Samuel
Hutter, Frank
author_facet Garg, Anurag
Ali, Muhammad
Hollmann, Noah
Purucker, Lennart
Müller, Samuel
Hutter, Frank
contents Foundation models for tabular data, like TabPFN, achieve strong performance on small datasets when pre-trained solely on synthetic data. We show that this performance can be significantly boosted by a targeted continued pre-training phase. Specifically, we demonstrate that leveraging a small, curated collection of large, real-world datasets for continued pre-training yields superior downstream predictive accuracy compared to using broader, potentially noisier corpora like CommonCrawl or GitTables. Our resulting model, Real-TabPFN, achieves substantial performance gains on 29 datasets from the OpenML AutoML Benchmark.
format Preprint
id arxiv_https___arxiv_org_abs_2507_03971
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Real-TabPFN: Improving Tabular Foundation Models via Continued Pre-training With Real-World Data
Garg, Anurag
Ali, Muhammad
Hollmann, Noah
Purucker, Lennart
Müller, Samuel
Hutter, Frank
Machine Learning
Artificial Intelligence
Methodology
Foundation models for tabular data, like TabPFN, achieve strong performance on small datasets when pre-trained solely on synthetic data. We show that this performance can be significantly boosted by a targeted continued pre-training phase. Specifically, we demonstrate that leveraging a small, curated collection of large, real-world datasets for continued pre-training yields superior downstream predictive accuracy compared to using broader, potentially noisier corpora like CommonCrawl or GitTables. Our resulting model, Real-TabPFN, achieves substantial performance gains on 29 datasets from the OpenML AutoML Benchmark.
title Real-TabPFN: Improving Tabular Foundation Models via Continued Pre-training With Real-World Data
topic Machine Learning
Artificial Intelligence
Methodology
url https://arxiv.org/abs/2507.03971