Data Presentation Over Architecture: Resampling Strategies for Credit Risk Prediction with Tabular Foundation Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tanna, Aditya, Solanki, Mitul, Bouadi, Mohamed, Bouarour, Nassim, Seth, Pratinav, Sankarapu, Vinay Kumar
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917508347330560
author Tanna, Aditya
Solanki, Mitul
Bouadi, Mohamed
Bouarour, Nassim
Seth, Pratinav
Sankarapu, Vinay Kumar
author_facet Tanna, Aditya
Solanki, Mitul
Bouadi, Mohamed
Bouarour, Nassim
Seth, Pratinav
Sankarapu, Vinay Kumar
contents Credit default prediction is a tabular learning problem with severe class imbalance, heterogeneous features, and tight latency budgets. Tabular Foundation Models (TFMs) approach this problem through in-context learning, which makes their predictions sensitive to how the context window is built. We benchmark four classical models and five TFMs on the Home Credit and Lending Club datasets, varying the context-construction strategy (seven options) and the context size (1K to 50K). On both datasets, the choice of context strategy explains more variance in AUC-ROC than the choice of TFM family: balanced and hybrid sampling add 3 to 4 AUC points over uniform sampling, and the gap exceeds the spread between TFMs. With a balanced context of 5K to 10K examples, the strongest TFMs reach the AUC of classical baselines trained on the full data, while also recovering meaningful default-class recall that default-threshold GBDTs do not. We frame this as evidence that context construction, rather than architecture choice, is the primary deployment lever for TFMs in imbalanced credit-risk settings.
format Preprint
id arxiv_https___arxiv_org_abs_2605_18635
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Data Presentation Over Architecture: Resampling Strategies for Credit Risk Prediction with Tabular Foundation Models
Tanna, Aditya
Solanki, Mitul
Bouadi, Mohamed
Bouarour, Nassim
Seth, Pratinav
Sankarapu, Vinay Kumar
Machine Learning
Artificial Intelligence
Credit default prediction is a tabular learning problem with severe class imbalance, heterogeneous features, and tight latency budgets. Tabular Foundation Models (TFMs) approach this problem through in-context learning, which makes their predictions sensitive to how the context window is built. We benchmark four classical models and five TFMs on the Home Credit and Lending Club datasets, varying the context-construction strategy (seven options) and the context size (1K to 50K). On both datasets, the choice of context strategy explains more variance in AUC-ROC than the choice of TFM family: balanced and hybrid sampling add 3 to 4 AUC points over uniform sampling, and the gap exceeds the spread between TFMs. With a balanced context of 5K to 10K examples, the strongest TFMs reach the AUC of classical baselines trained on the full data, while also recovering meaningful default-class recall that default-threshold GBDTs do not. We frame this as evidence that context construction, rather than architecture choice, is the primary deployment lever for TFMs in imbalanced credit-risk settings.
title Data Presentation Over Architecture: Resampling Strategies for Credit Risk Prediction with Tabular Foundation Models
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2605.18635