Towards Universal Tabular Embeddings: A Benchmark Across Data Tasks

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Vogel, Liane, Srinivas, Kavitha, D'Souza, Niharika, Shirai, Sola, Hassanzadeh, Oktie, Samulowitz, Horst
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915952160931840
author Vogel, Liane
Srinivas, Kavitha
D'Souza, Niharika
Shirai, Sola
Hassanzadeh, Oktie
Samulowitz, Horst
author_facet Vogel, Liane
Srinivas, Kavitha
D'Souza, Niharika
Shirai, Sola
Hassanzadeh, Oktie
Samulowitz, Horst
contents Tabular foundation models aim to learn universal representations of tabular data that transfer across tasks and domains, enabling applications such as table retrieval, semantic search and table-based prediction. Despite the growing number of such models, it remains unclear which approach works best in practice, as existing methods are often evaluated under task-specific settings that make direct comparison difficult. To address this, we introduce TEmBed, the Tabular Embedding Test Bed, a comprehensive benchmark for systematically evaluating tabular embeddings across four representation levels: cell, row, column, and table. Evaluating a diverse set of tabular representation learning models, we show that which model to use depends on the task and representation level. Our results offer practical guidance for selecting tabular embeddings in real-world applications and lay the groundwork for developing more general-purpose tabular representation models.
format Preprint
id arxiv_https___arxiv_org_abs_2604_21696
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Towards Universal Tabular Embeddings: A Benchmark Across Data Tasks
Vogel, Liane
Srinivas, Kavitha
D'Souza, Niharika
Shirai, Sola
Hassanzadeh, Oktie
Samulowitz, Horst
Machine Learning
Databases
Tabular foundation models aim to learn universal representations of tabular data that transfer across tasks and domains, enabling applications such as table retrieval, semantic search and table-based prediction. Despite the growing number of such models, it remains unclear which approach works best in practice, as existing methods are often evaluated under task-specific settings that make direct comparison difficult. To address this, we introduce TEmBed, the Tabular Embedding Test Bed, a comprehensive benchmark for systematically evaluating tabular embeddings across four representation levels: cell, row, column, and table. Evaluating a diverse set of tabular representation learning models, we show that which model to use depends on the task and representation level. Our results offer practical guidance for selecting tabular embeddings in real-world applications and lay the groundwork for developing more general-purpose tabular representation models.
title Towards Universal Tabular Embeddings: A Benchmark Across Data Tasks
topic Machine Learning
Databases
url https://arxiv.org/abs/2604.21696