MulTaBench: Benchmarking Multimodal Tabular Learning with Text and Image

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Arazi, Alan, Shapira, Eilam, Grunblat, Shoham, Ventura, Mor, Hoffer, Elad, Blayer, Gioia, Holzmüller, David, Purucker, Lennart, Varoquaux, Gaël, Hutter, Frank, Reichart, Roi
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911671704879104
author Arazi, Alan
Shapira, Eilam
Grunblat, Shoham
Ventura, Mor
Hoffer, Elad
Blayer, Gioia
Holzmüller, David
Purucker, Lennart
Varoquaux, Gaël
Hutter, Frank
Reichart, Roi
author_facet Arazi, Alan
Shapira, Eilam
Grunblat, Shoham
Ventura, Mor
Hoffer, Elad
Blayer, Gioia
Holzmüller, David
Purucker, Lennart
Varoquaux, Gaël
Hutter, Frank
Reichart, Roi
contents Tabular Foundation Models have recently established the state of the art in supervised tabular learning, by leveraging pretraining to learn generalizable representations of numerical and categorical structured data. However, they lack native support for unstructured modalities such as text and image, and rely on frozen, pretrained embeddings to process them. On established Multimodal Tabular Learning benchmarks, we show that tuning the embeddings to the task improves performance. Existing benchmarks, however, often focus on the mere co-occurrence of modalities; this leads to high variance across datasets and masks the benefits of task-specific tuning. To address this gap, we introduce MulTaBench, a benchmark of 40 datasets, split equally between image-tabular and text-tabular tasks. We focus on predictive tasks where the modalities provide complementary predictive signal, and where generic embeddings lose critical information, necessitating Target-Aware Representations that are aligned with the task. Our experimental results demonstrate that the gains from target-aware representation tuning generalize across both text and image modalities, several tabular learners, encoder scales, and embedding dimensions. MulTaBench constitutes the largest image-tabular benchmarking effort to date, spanning high-impact domains such as healthcare and e-commerce. It is designed to enable the research of novel architectures which incorporate joint modeling and target-aware representations, paving the way for the development of novel Multimodal Tabular Foundation Models.
format Preprint
id arxiv_https___arxiv_org_abs_2605_10616
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle MulTaBench: Benchmarking Multimodal Tabular Learning with Text and Image
Arazi, Alan
Shapira, Eilam
Grunblat, Shoham
Ventura, Mor
Hoffer, Elad
Blayer, Gioia
Holzmüller, David
Purucker, Lennart
Varoquaux, Gaël
Hutter, Frank
Reichart, Roi
Machine Learning
Computation and Language
Computer Vision and Pattern Recognition
Tabular Foundation Models have recently established the state of the art in supervised tabular learning, by leveraging pretraining to learn generalizable representations of numerical and categorical structured data. However, they lack native support for unstructured modalities such as text and image, and rely on frozen, pretrained embeddings to process them. On established Multimodal Tabular Learning benchmarks, we show that tuning the embeddings to the task improves performance. Existing benchmarks, however, often focus on the mere co-occurrence of modalities; this leads to high variance across datasets and masks the benefits of task-specific tuning. To address this gap, we introduce MulTaBench, a benchmark of 40 datasets, split equally between image-tabular and text-tabular tasks. We focus on predictive tasks where the modalities provide complementary predictive signal, and where generic embeddings lose critical information, necessitating Target-Aware Representations that are aligned with the task. Our experimental results demonstrate that the gains from target-aware representation tuning generalize across both text and image modalities, several tabular learners, encoder scales, and embedding dimensions. MulTaBench constitutes the largest image-tabular benchmarking effort to date, spanning high-impact domains such as healthcare and e-commerce. It is designed to enable the research of novel architectures which incorporate joint modeling and target-aware representations, paving the way for the development of novel Multimodal Tabular Foundation Models.
title MulTaBench: Benchmarking Multimodal Tabular Learning with Text and Image
topic Machine Learning
Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.10616