Towards Universal Tabular Embeddings: A Benchmark Across Data Tasks
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866915952160931840 |
|---|---|
| author | Vogel, Liane Srinivas, Kavitha D'Souza, Niharika Shirai, Sola Hassanzadeh, Oktie Samulowitz, Horst |
| author_facet | Vogel, Liane Srinivas, Kavitha D'Souza, Niharika Shirai, Sola Hassanzadeh, Oktie Samulowitz, Horst |
| contents | Tabular foundation models aim to learn universal representations of tabular data that transfer across tasks and domains, enabling applications such as table retrieval, semantic search and table-based prediction. Despite the growing number of such models, it remains unclear which approach works best in practice, as existing methods are often evaluated under task-specific settings that make direct comparison difficult. To address this, we introduce TEmBed, the Tabular Embedding Test Bed, a comprehensive benchmark for systematically evaluating tabular embeddings across four representation levels: cell, row, column, and table. Evaluating a diverse set of tabular representation learning models, we show that which model to use depends on the task and representation level. Our results offer practical guidance for selecting tabular embeddings in real-world applications and lay the groundwork for developing more general-purpose tabular representation models. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2604_21696 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Towards Universal Tabular Embeddings: A Benchmark Across Data Tasks Vogel, Liane Srinivas, Kavitha D'Souza, Niharika Shirai, Sola Hassanzadeh, Oktie Samulowitz, Horst Machine Learning Databases Tabular foundation models aim to learn universal representations of tabular data that transfer across tasks and domains, enabling applications such as table retrieval, semantic search and table-based prediction. Despite the growing number of such models, it remains unclear which approach works best in practice, as existing methods are often evaluated under task-specific settings that make direct comparison difficult. To address this, we introduce TEmBed, the Tabular Embedding Test Bed, a comprehensive benchmark for systematically evaluating tabular embeddings across four representation levels: cell, row, column, and table. Evaluating a diverse set of tabular representation learning models, we show that which model to use depends on the task and representation level. Our results offer practical guidance for selecting tabular embeddings in real-world applications and lay the groundwork for developing more general-purpose tabular representation models. |
| title | Towards Universal Tabular Embeddings: A Benchmark Across Data Tasks |
| topic | Machine Learning Databases |
| url | https://arxiv.org/abs/2604.21696 |