SPINACH: SPARQL-Based Information Navigation for Challenging Real-World Questions

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Liu, Shicheng, Semnani, Sina J., Triedman, Harold, Xu, Jialiang, Zhao, Isaac Dan, Lam, Monica S.
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866929551146221568
author Liu, Shicheng
Semnani, Sina J.
Triedman, Harold
Xu, Jialiang
Zhao, Isaac Dan
Lam, Monica S.
author_facet Liu, Shicheng
Semnani, Sina J.
Triedman, Harold
Xu, Jialiang
Zhao, Isaac Dan
Lam, Monica S.
contents Large Language Models (LLMs) have led to significant improvements in the Knowledge Base Question Answering (KBQA) task. However, datasets used in KBQA studies do not capture the true complexity of KBQA tasks. They either have simple questions, use synthetically generated logical forms, or are based on small knowledge base (KB) schemas. We introduce the SPINACH dataset, an expert-annotated KBQA dataset collected from discussions on Wikidata's "Request a Query" forum with 320 decontextualized question-SPARQL pairs. The complexity of these in-the-wild queries calls for a KBQA system that can dynamically explore large and often incomplete schemas and reason about them, as it is infeasible to create a comprehensive training dataset. We also introduce an in-context learning KBQA agent, also called SPINACH, that mimics how a human expert would write SPARQLs to handle challenging questions. SPINACH achieves a new state of the art on the QALD-7, QALD-9 Plus and QALD-10 datasets by 31.0%, 27.0%, and 10.0% in $F_1$, respectively, and coming within 1.6% of the fine-tuned LLaMA SOTA model on WikiWebQuestions. On our new SPINACH dataset, the SPINACH agent outperforms all baselines, including the best GPT-4-based KBQA agent, by at least 38.1% in $F_1$.
format Preprint
id arxiv_https___arxiv_org_abs_2407_11417
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle SPINACH: SPARQL-Based Information Navigation for Challenging Real-World Questions
Liu, Shicheng
Semnani, Sina J.
Triedman, Harold
Xu, Jialiang
Zhao, Isaac Dan
Lam, Monica S.
Computation and Language
Large Language Models (LLMs) have led to significant improvements in the Knowledge Base Question Answering (KBQA) task. However, datasets used in KBQA studies do not capture the true complexity of KBQA tasks. They either have simple questions, use synthetically generated logical forms, or are based on small knowledge base (KB) schemas. We introduce the SPINACH dataset, an expert-annotated KBQA dataset collected from discussions on Wikidata's "Request a Query" forum with 320 decontextualized question-SPARQL pairs. The complexity of these in-the-wild queries calls for a KBQA system that can dynamically explore large and often incomplete schemas and reason about them, as it is infeasible to create a comprehensive training dataset. We also introduce an in-context learning KBQA agent, also called SPINACH, that mimics how a human expert would write SPARQLs to handle challenging questions. SPINACH achieves a new state of the art on the QALD-7, QALD-9 Plus and QALD-10 datasets by 31.0%, 27.0%, and 10.0% in $F_1$, respectively, and coming within 1.6% of the fine-tuned LLaMA SOTA model on WikiWebQuestions. On our new SPINACH dataset, the SPINACH agent outperforms all baselines, including the best GPT-4-based KBQA agent, by at least 38.1% in $F_1$.
title SPINACH: SPARQL-Based Information Navigation for Challenging Real-World Questions
topic Computation and Language
url https://arxiv.org/abs/2407.11417