TabEmbed: Benchmarking and Learning Generalist Embeddings for Tabular Understanding

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Qiang, Minjie, Zhang, Mingming, Bao, Xiaoyi, Fu, Xing, Cheng, Yu, Wang, Weiqiang, Wang, Zhongqing, Wang, Ningtao
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913094820691968
author Qiang, Minjie
Zhang, Mingming
Bao, Xiaoyi
Fu, Xing
Cheng, Yu
Wang, Weiqiang
Wang, Zhongqing
Wang, Ningtao
author_facet Qiang, Minjie
Zhang, Mingming
Bao, Xiaoyi
Fu, Xing
Cheng, Yu
Wang, Weiqiang
Wang, Zhongqing
Wang, Ningtao
contents Foundation models have established unified representations for natural language processing, yet this paradigm remains largely unexplored for tabular data. Existing methods face fundamental limitations: LLM-based approaches lack retrieval-compatible vector outputs, whereas text embedding models often fail to capture tabular structure and numerical semantics. To bridge this gap, we first introduce the Tabular Embedding Benchmark (TabBench), a comprehensive suite designed to evaluate the tabular understanding capability of embedding models. We then propose TabEmbed, the first generalist embedding model that unifies tabular classification and retrieval within a shared embedding space. By reformulating diverse tabular tasks as semantic matching problems, TabEmbed leverages large-scale contrastive learning with positive-aware hard negative mining to discern fine-grained structural and numerical nuances. Experimental results on TabBench demonstrate that TabEmbed significantly outperforms state-of-the-art text embedding models, establishing a new baseline for universal tabular representation learning. Code and datasets are publicly available at https://github.com/qiangminjie27/TabEmbed and https://huggingface.co/datasets/qiangminjie27/TabBench.
format Preprint
id arxiv_https___arxiv_org_abs_2605_04962
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle TabEmbed: Benchmarking and Learning Generalist Embeddings for Tabular Understanding
Qiang, Minjie
Zhang, Mingming
Bao, Xiaoyi
Fu, Xing
Cheng, Yu
Wang, Weiqiang
Wang, Zhongqing
Wang, Ningtao
Computation and Language
Information Retrieval
Foundation models have established unified representations for natural language processing, yet this paradigm remains largely unexplored for tabular data. Existing methods face fundamental limitations: LLM-based approaches lack retrieval-compatible vector outputs, whereas text embedding models often fail to capture tabular structure and numerical semantics. To bridge this gap, we first introduce the Tabular Embedding Benchmark (TabBench), a comprehensive suite designed to evaluate the tabular understanding capability of embedding models. We then propose TabEmbed, the first generalist embedding model that unifies tabular classification and retrieval within a shared embedding space. By reformulating diverse tabular tasks as semantic matching problems, TabEmbed leverages large-scale contrastive learning with positive-aware hard negative mining to discern fine-grained structural and numerical nuances. Experimental results on TabBench demonstrate that TabEmbed significantly outperforms state-of-the-art text embedding models, establishing a new baseline for universal tabular representation learning. Code and datasets are publicly available at https://github.com/qiangminjie27/TabEmbed and https://huggingface.co/datasets/qiangminjie27/TabBench.
title TabEmbed: Benchmarking and Learning Generalist Embeddings for Tabular Understanding
topic Computation and Language
Information Retrieval
url https://arxiv.org/abs/2605.04962