Real-TabPFN: Improving Tabular Foundation Models via Continued Pre-training With Real-World Data
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909676012044288 |
|---|---|
| author | Garg, Anurag Ali, Muhammad Hollmann, Noah Purucker, Lennart Müller, Samuel Hutter, Frank |
| author_facet | Garg, Anurag Ali, Muhammad Hollmann, Noah Purucker, Lennart Müller, Samuel Hutter, Frank |
| contents | Foundation models for tabular data, like TabPFN, achieve strong performance on small datasets when pre-trained solely on synthetic data. We show that this performance can be significantly boosted by a targeted continued pre-training phase. Specifically, we demonstrate that leveraging a small, curated collection of large, real-world datasets for continued pre-training yields superior downstream predictive accuracy compared to using broader, potentially noisier corpora like CommonCrawl or GitTables. Our resulting model, Real-TabPFN, achieves substantial performance gains on 29 datasets from the OpenML AutoML Benchmark. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2507_03971 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Real-TabPFN: Improving Tabular Foundation Models via Continued Pre-training With Real-World Data Garg, Anurag Ali, Muhammad Hollmann, Noah Purucker, Lennart Müller, Samuel Hutter, Frank Machine Learning Artificial Intelligence Methodology Foundation models for tabular data, like TabPFN, achieve strong performance on small datasets when pre-trained solely on synthetic data. We show that this performance can be significantly boosted by a targeted continued pre-training phase. Specifically, we demonstrate that leveraging a small, curated collection of large, real-world datasets for continued pre-training yields superior downstream predictive accuracy compared to using broader, potentially noisier corpora like CommonCrawl or GitTables. Our resulting model, Real-TabPFN, achieves substantial performance gains on 29 datasets from the OpenML AutoML Benchmark. |
| title | Real-TabPFN: Improving Tabular Foundation Models via Continued Pre-training With Real-World Data |
| topic | Machine Learning Artificial Intelligence Methodology |
| url | https://arxiv.org/abs/2507.03971 |