TabEmbed: Benchmarking and Learning Generalist Embeddings for Tabular Understanding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866913094820691968 |
|---|---|
| author | Qiang, Minjie Zhang, Mingming Bao, Xiaoyi Fu, Xing Cheng, Yu Wang, Weiqiang Wang, Zhongqing Wang, Ningtao |
| author_facet | Qiang, Minjie Zhang, Mingming Bao, Xiaoyi Fu, Xing Cheng, Yu Wang, Weiqiang Wang, Zhongqing Wang, Ningtao |
| contents | Foundation models have established unified representations for natural language processing, yet this paradigm remains largely unexplored for tabular data. Existing methods face fundamental limitations: LLM-based approaches lack retrieval-compatible vector outputs, whereas text embedding models often fail to capture tabular structure and numerical semantics. To bridge this gap, we first introduce the Tabular Embedding Benchmark (TabBench), a comprehensive suite designed to evaluate the tabular understanding capability of embedding models. We then propose TabEmbed, the first generalist embedding model that unifies tabular classification and retrieval within a shared embedding space. By reformulating diverse tabular tasks as semantic matching problems, TabEmbed leverages large-scale contrastive learning with positive-aware hard negative mining to discern fine-grained structural and numerical nuances. Experimental results on TabBench demonstrate that TabEmbed significantly outperforms state-of-the-art text embedding models, establishing a new baseline for universal tabular representation learning. Code and datasets are publicly available at https://github.com/qiangminjie27/TabEmbed and https://huggingface.co/datasets/qiangminjie27/TabBench. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2605_04962 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | TabEmbed: Benchmarking and Learning Generalist Embeddings for Tabular Understanding Qiang, Minjie Zhang, Mingming Bao, Xiaoyi Fu, Xing Cheng, Yu Wang, Weiqiang Wang, Zhongqing Wang, Ningtao Computation and Language Information Retrieval Foundation models have established unified representations for natural language processing, yet this paradigm remains largely unexplored for tabular data. Existing methods face fundamental limitations: LLM-based approaches lack retrieval-compatible vector outputs, whereas text embedding models often fail to capture tabular structure and numerical semantics. To bridge this gap, we first introduce the Tabular Embedding Benchmark (TabBench), a comprehensive suite designed to evaluate the tabular understanding capability of embedding models. We then propose TabEmbed, the first generalist embedding model that unifies tabular classification and retrieval within a shared embedding space. By reformulating diverse tabular tasks as semantic matching problems, TabEmbed leverages large-scale contrastive learning with positive-aware hard negative mining to discern fine-grained structural and numerical nuances. Experimental results on TabBench demonstrate that TabEmbed significantly outperforms state-of-the-art text embedding models, establishing a new baseline for universal tabular representation learning. Code and datasets are publicly available at https://github.com/qiangminjie27/TabEmbed and https://huggingface.co/datasets/qiangminjie27/TabBench. |
| title | TabEmbed: Benchmarking and Learning Generalist Embeddings for Tabular Understanding |
| topic | Computation and Language Information Retrieval |
| url | https://arxiv.org/abs/2605.04962 |