nanoTabPFN: A Lightweight and Educational Reimplementation of TabPFN

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Pfefferle, Alexander, Hog, Johannes, Purucker, Lennart, Hutter, Frank
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918253723385856
author Pfefferle, Alexander
Hog, Johannes
Purucker, Lennart
Hutter, Frank
author_facet Pfefferle, Alexander
Hog, Johannes
Purucker, Lennart
Hutter, Frank
contents Tabular foundation models such as TabPFN have revolutionized predictive machine learning for tabular data. At the same time, the driving factors of this revolution are hard to understand. Existing open-source tabular foundation models are implemented in complicated pipelines boasting over 10,000 lines of code, lack architecture documentation or code quality. In short, the implementations are hard to understand, not beginner-friendly, and complicated to adapt for new experiments. We introduce nanoTabPFN, a simplified and lightweight implementation of the TabPFN v2 architecture and a corresponding training loop that uses pre-generated training data. nanoTabPFN makes tabular foundation models more accessible to students and researchers alike. For example, restricted to a small data setting it achieves a performance comparable to traditional machine learning baselines within one minute of pre-training on a single GPU (160,000x faster than TabPFN v2 pretraining). This eliminated requirement of large computational resources makes pre-training tabular foundation models accessible for educational purposes. Our code is available at https://github.com/automl/nanoTabPFN.
format Preprint
id arxiv_https___arxiv_org_abs_2511_03634
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle nanoTabPFN: A Lightweight and Educational Reimplementation of TabPFN
Pfefferle, Alexander
Hog, Johannes
Purucker, Lennart
Hutter, Frank
Machine Learning
Tabular foundation models such as TabPFN have revolutionized predictive machine learning for tabular data. At the same time, the driving factors of this revolution are hard to understand. Existing open-source tabular foundation models are implemented in complicated pipelines boasting over 10,000 lines of code, lack architecture documentation or code quality. In short, the implementations are hard to understand, not beginner-friendly, and complicated to adapt for new experiments. We introduce nanoTabPFN, a simplified and lightweight implementation of the TabPFN v2 architecture and a corresponding training loop that uses pre-generated training data. nanoTabPFN makes tabular foundation models more accessible to students and researchers alike. For example, restricted to a small data setting it achieves a performance comparable to traditional machine learning baselines within one minute of pre-training on a single GPU (160,000x faster than TabPFN v2 pretraining). This eliminated requirement of large computational resources makes pre-training tabular foundation models accessible for educational purposes. Our code is available at https://github.com/automl/nanoTabPFN.
title nanoTabPFN: A Lightweight and Educational Reimplementation of TabPFN
topic Machine Learning
url https://arxiv.org/abs/2511.03634