Hybrid-SQuAD: Hybrid Scholarly Question Answering Dataset

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Taffa, Tilahun Abedissa, Banerjee, Debayan, Assabie, Yaregal, Usbeck, Ricardo
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912144860119040
author Taffa, Tilahun Abedissa
Banerjee, Debayan
Assabie, Yaregal
Usbeck, Ricardo
author_facet Taffa, Tilahun Abedissa
Banerjee, Debayan
Assabie, Yaregal
Usbeck, Ricardo
contents Existing Scholarly Question Answering (QA) methods typically target homogeneous data sources, relying solely on either text or Knowledge Graphs (KGs). However, scholarly information often spans heterogeneous sources, necessitating the development of QA systems that integrate information from multiple heterogeneous data sources. To address this challenge, we introduce Hybrid-SQuAD (Hybrid Scholarly Question Answering Dataset), a novel large-scale QA dataset designed to facilitate answering questions incorporating both text and KG facts. The dataset consists of 10.5K question-answer pairs generated by a large language model, leveraging the KGs DBLP and SemOpenAlex alongside corresponding text from Wikipedia. In addition, we propose a RAG-based baseline hybrid QA model, achieving an exact match score of 69.65 on the Hybrid-SQuAD test set.
format Preprint
id arxiv_https___arxiv_org_abs_2412_02788
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Hybrid-SQuAD: Hybrid Scholarly Question Answering Dataset
Taffa, Tilahun Abedissa
Banerjee, Debayan
Assabie, Yaregal
Usbeck, Ricardo
Computation and Language
Artificial Intelligence
Existing Scholarly Question Answering (QA) methods typically target homogeneous data sources, relying solely on either text or Knowledge Graphs (KGs). However, scholarly information often spans heterogeneous sources, necessitating the development of QA systems that integrate information from multiple heterogeneous data sources. To address this challenge, we introduce Hybrid-SQuAD (Hybrid Scholarly Question Answering Dataset), a novel large-scale QA dataset designed to facilitate answering questions incorporating both text and KG facts. The dataset consists of 10.5K question-answer pairs generated by a large language model, leveraging the KGs DBLP and SemOpenAlex alongside corresponding text from Wikipedia. In addition, we propose a RAG-based baseline hybrid QA model, achieving an exact match score of 69.65 on the Hybrid-SQuAD test set.
title Hybrid-SQuAD: Hybrid Scholarly Question Answering Dataset
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2412.02788