DPDisc: From Factoid Questions to Data Product Requests for Open-World Data Product Discovery over Tables and Text

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Liangliang, Mihindukulasooriya, Nandana, D'Souza, Niharika S., Shirai, Sola, Dash, Sarthak, Ma, Yao, Samulowitz, Horst
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914406639599616
author Zhang, Liangliang
Mihindukulasooriya, Nandana
D'Souza, Niharika S.
Shirai, Sola
Dash, Sarthak
Ma, Yao
Samulowitz, Horst
author_facet Zhang, Liangliang
Mihindukulasooriya, Nandana
D'Souza, Niharika S.
Shirai, Sola
Dash, Sarthak
Ma, Yao
Samulowitz, Horst
contents Data products are reusable, self-contained assets designed for specific business use cases. Automating their discovery is of great industry interest, as it enables efficient data access in large data lakes and supports analytical workflows. However, no benchmark currently exists for data product discovery over hybrid table-text corpora. Existing datasets focus on answering single factoid questions over individual tables rather than assembling multiple related data assets into coherent products. To address this gap, we present DPDisc, the first large-scale benchmark for data product discovery, where systems must retrieve coherent collections of tables and passages to satisfy high-level Data Product Requests (DPRs). We introduce DPForge, an automated pipeline that systematically repurposes table-text QA datasets by clustering related tables and passages into coherent data products, generating professional-level analytical requests using an LLM ensemble, and validating quality through multi-phase LLM evaluation. DPDisc comprises 13,076 validated instances with full provenance, derived from three representative datasets spanning open-domain and financial domains. Baseline experiments with sparse, dense, and hybrid retrieval methods imply evaluation feasibility while revealing substantial performance gaps across domains, indicating opportunities for future research in structure-aware data product discovery. Code and datasets are available at: Dataset: https://huggingface.co/datasets/ibm-research/data-product-benchmark Code: https://github.com/ibm/data-product-benchmark
format Preprint
id arxiv_https___arxiv_org_abs_2510_21737
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DPDisc: From Factoid Questions to Data Product Requests for Open-World Data Product Discovery over Tables and Text
Zhang, Liangliang
Mihindukulasooriya, Nandana
D'Souza, Niharika S.
Shirai, Sola
Dash, Sarthak
Ma, Yao
Samulowitz, Horst
Information Retrieval
68T30, 68T50
I.2.7; I.2.4; H.3.3
Data products are reusable, self-contained assets designed for specific business use cases. Automating their discovery is of great industry interest, as it enables efficient data access in large data lakes and supports analytical workflows. However, no benchmark currently exists for data product discovery over hybrid table-text corpora. Existing datasets focus on answering single factoid questions over individual tables rather than assembling multiple related data assets into coherent products. To address this gap, we present DPDisc, the first large-scale benchmark for data product discovery, where systems must retrieve coherent collections of tables and passages to satisfy high-level Data Product Requests (DPRs). We introduce DPForge, an automated pipeline that systematically repurposes table-text QA datasets by clustering related tables and passages into coherent data products, generating professional-level analytical requests using an LLM ensemble, and validating quality through multi-phase LLM evaluation. DPDisc comprises 13,076 validated instances with full provenance, derived from three representative datasets spanning open-domain and financial domains. Baseline experiments with sparse, dense, and hybrid retrieval methods imply evaluation feasibility while revealing substantial performance gaps across domains, indicating opportunities for future research in structure-aware data product discovery. Code and datasets are available at: Dataset: https://huggingface.co/datasets/ibm-research/data-product-benchmark Code: https://github.com/ibm/data-product-benchmark
title DPDisc: From Factoid Questions to Data Product Requests for Open-World Data Product Discovery over Tables and Text
topic Information Retrieval
68T30, 68T50
I.2.7; I.2.4; H.3.3
url https://arxiv.org/abs/2510.21737