Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Shi, Qing, He, Jing, Chen, Qiaosheng, Cheng, Gong
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:https://arxiv.org/abs/2510.17228
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914190551154688
author Shi, Qing
He, Jing
Chen, Qiaosheng
Cheng, Gong
author_facet Shi, Qing
He, Jing
Chen, Qiaosheng
Cheng, Gong
contents Dataset search is a well-established task in the Semantic Web and information retrieval research. Current approaches retrieve datasets either based on keyword queries or by identifying datasets similar to a given target dataset. These paradigms fail when the information need involves both keywords and target datasets. To address this gap, we investigate a generalized task, Dataset Search with Examples (DSE), and extend it to Explainable DSE (ExDSE), which further requires identifying relevant fields of the retrieved datasets. We construct DSEBench, the first test collection that provides high-quality dataset-level and field-level annotations to support the evaluation of DSE and ExDSE, respectively. In addition, we employ a large language model to generate extensive annotations for training purposes. We establish comprehensive baselines on DSEBench by adapting and evaluating a variety of lexical, dense, and LLM-based retrieval, reranking, and explanation methods.
format Preprint
id arxiv_https___arxiv_org_abs_2510_17228
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DSEBench: A Test Collection for Explainable Dataset Search with Examples
Shi, Qing
He, Jing
Chen, Qiaosheng
Cheng, Gong
Information Retrieval
Dataset search is a well-established task in the Semantic Web and information retrieval research. Current approaches retrieve datasets either based on keyword queries or by identifying datasets similar to a given target dataset. These paradigms fail when the information need involves both keywords and target datasets. To address this gap, we investigate a generalized task, Dataset Search with Examples (DSE), and extend it to Explainable DSE (ExDSE), which further requires identifying relevant fields of the retrieved datasets. We construct DSEBench, the first test collection that provides high-quality dataset-level and field-level annotations to support the evaluation of DSE and ExDSE, respectively. In addition, we employ a large language model to generate extensive annotations for training purposes. We establish comprehensive baselines on DSEBench by adapting and evaluating a variety of lexical, dense, and LLM-based retrieval, reranking, and explanation methods.
title DSEBench: A Test Collection for Explainable Dataset Search with Examples
topic Information Retrieval
url https://arxiv.org/abs/2510.17228