Something's Fishy In The Data Lake: A Critical Re-evaluation of Table Union Search Benchmarks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Boutaleb, Allaa, Amann, Bernd, Naacke, Hubert, Angarita, Rafael
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908382638637056
author Boutaleb, Allaa
Amann, Bernd
Naacke, Hubert
Angarita, Rafael
author_facet Boutaleb, Allaa
Amann, Bernd
Naacke, Hubert
Angarita, Rafael
contents Recent table representation learning and data discovery methods tackle table union search (TUS) within data lakes, which involves identifying tables that can be unioned with a given query table to enrich its content. These methods are commonly evaluated using benchmarks that aim to assess semantic understanding in real-world TUS tasks. However, our analysis of prominent TUS benchmarks reveals several limitations that allow simple baselines to perform surprisingly well, often outperforming more sophisticated approaches. This suggests that current benchmark scores are heavily influenced by dataset-specific characteristics and fail to effectively isolate the gains from semantic understanding. To address this, we propose essential criteria for future benchmarks to enable a more realistic and reliable evaluation of progress in semantic table union search.
format Preprint
id arxiv_https___arxiv_org_abs_2505_21329
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Something's Fishy In The Data Lake: A Critical Re-evaluation of Table Union Search Benchmarks
Boutaleb, Allaa
Amann, Bernd
Naacke, Hubert
Angarita, Rafael
Information Retrieval
Artificial Intelligence
Computation and Language
Databases
Machine Learning
Recent table representation learning and data discovery methods tackle table union search (TUS) within data lakes, which involves identifying tables that can be unioned with a given query table to enrich its content. These methods are commonly evaluated using benchmarks that aim to assess semantic understanding in real-world TUS tasks. However, our analysis of prominent TUS benchmarks reveals several limitations that allow simple baselines to perform surprisingly well, often outperforming more sophisticated approaches. This suggests that current benchmark scores are heavily influenced by dataset-specific characteristics and fail to effectively isolate the gains from semantic understanding. To address this, we propose essential criteria for future benchmarks to enable a more realistic and reliable evaluation of progress in semantic table union search.
title Something's Fishy In The Data Lake: A Critical Re-evaluation of Table Union Search Benchmarks
topic Information Retrieval
Artificial Intelligence
Computation and Language
Databases
Machine Learning
url https://arxiv.org/abs/2505.21329