NLCTables: A Dataset for Marrying Natural Language Conditions with Table Discovery

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cui, Lingxi, Li, Huan, Chen, Ke, Shou, Lidan, Chen, Gang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910916500520960
author Cui, Lingxi
Li, Huan
Chen, Ke
Shou, Lidan
Chen, Gang
author_facet Cui, Lingxi
Li, Huan
Chen, Ke
Shou, Lidan
Chen, Gang
contents With the growing abundance of repositories containing tabular data, discovering relevant tables for in-depth analysis remains a challenging task. Existing table discovery methods primarily retrieve desired tables based on a query table or several vague keywords, leaving users to manually filter large result sets. To address this limitation, we propose a new task: NL-conditional table discovery (nlcTD), where users combine a query table with natural language (NL) requirements to refine search results. To advance research in this area, we present nlcTables, a comprehensive benchmark dataset comprising 627 diverse queries spanning NL-only, union, join, and fuzzy conditions, 22,080 candidate tables, and 21,200 relevance annotations. Our evaluation of six state-of-the-art table discovery methods on nlcTables reveals substantial performance gaps, highlighting the need for advanced techniques to tackle this challenging nlcTD scenario. The dataset, construction framework, and baseline implementations are publicly available at https://github.com/SuDIS-ZJU/nlcTables to foster future research.
format Preprint
id arxiv_https___arxiv_org_abs_2504_15849
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle NLCTables: A Dataset for Marrying Natural Language Conditions with Table Discovery
Cui, Lingxi
Li, Huan
Chen, Ke
Shou, Lidan
Chen, Gang
Information Retrieval
68P20
With the growing abundance of repositories containing tabular data, discovering relevant tables for in-depth analysis remains a challenging task. Existing table discovery methods primarily retrieve desired tables based on a query table or several vague keywords, leaving users to manually filter large result sets. To address this limitation, we propose a new task: NL-conditional table discovery (nlcTD), where users combine a query table with natural language (NL) requirements to refine search results. To advance research in this area, we present nlcTables, a comprehensive benchmark dataset comprising 627 diverse queries spanning NL-only, union, join, and fuzzy conditions, 22,080 candidate tables, and 21,200 relevance annotations. Our evaluation of six state-of-the-art table discovery methods on nlcTables reveals substantial performance gaps, highlighting the need for advanced techniques to tackle this challenging nlcTD scenario. The dataset, construction framework, and baseline implementations are publicly available at https://github.com/SuDIS-ZJU/nlcTables to foster future research.
title NLCTables: A Dataset for Marrying Natural Language Conditions with Table Discovery
topic Information Retrieval
68P20
url https://arxiv.org/abs/2504.15849