TabDistill: Distilling Transformers into Neural Nets for Few-Shot Tabular Classification

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dissanayake, Pasan, Dutta, Sanghamitra
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909915330641920
author Dissanayake, Pasan
Dutta, Sanghamitra
author_facet Dissanayake, Pasan
Dutta, Sanghamitra
contents Transformer-based models have shown promising performance on tabular data compared to their classical counterparts such as neural networks and Gradient Boosted Decision Trees (GBDTs) in scenarios with limited training data. They utilize their pre-trained knowledge to adapt to new domains, achieving commendable performance with only a few training examples, also called the few-shot regime. However, the performance gain in the few-shot regime comes at the expense of significantly increased complexity and number of parameters. To circumvent this trade-off, we introduce TabDistill, a new strategy to distill the pre-trained knowledge in complex transformer-based models into simpler neural networks for effectively classifying tabular data. Our framework yields the best of both worlds: being parameter-efficient while performing well with limited training data. The distilled neural networks surpass classical baselines such as regular neural networks, XGBoost and logistic regression under equal training data, and in some cases, even the original transformer-based models that they were distilled from.
format Preprint
id arxiv_https___arxiv_org_abs_2511_05704
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle TabDistill: Distilling Transformers into Neural Nets for Few-Shot Tabular Classification
Dissanayake, Pasan
Dutta, Sanghamitra
Machine Learning
Artificial Intelligence
Computation and Language
Transformer-based models have shown promising performance on tabular data compared to their classical counterparts such as neural networks and Gradient Boosted Decision Trees (GBDTs) in scenarios with limited training data. They utilize their pre-trained knowledge to adapt to new domains, achieving commendable performance with only a few training examples, also called the few-shot regime. However, the performance gain in the few-shot regime comes at the expense of significantly increased complexity and number of parameters. To circumvent this trade-off, we introduce TabDistill, a new strategy to distill the pre-trained knowledge in complex transformer-based models into simpler neural networks for effectively classifying tabular data. Our framework yields the best of both worlds: being parameter-efficient while performing well with limited training data. The distilled neural networks surpass classical baselines such as regular neural networks, XGBoost and logistic regression under equal training data, and in some cases, even the original transformer-based models that they were distilled from.
title TabDistill: Distilling Transformers into Neural Nets for Few-Shot Tabular Classification
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2511.05704