Data Distribution Bottlenecks in Grounding Language Models to Knowledge Bases

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shu, Yiheng, Yu, Zhiwei
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916118656974848
author Shu, Yiheng
Yu, Zhiwei
author_facet Shu, Yiheng
Yu, Zhiwei
contents Language models (LMs) have already demonstrated remarkable abilities in understanding and generating both natural and formal language. Despite these advances, their integration with real-world environments such as large-scale knowledge bases (KBs) remains an underdeveloped area, affecting applications such as semantic parsing and indulging in "hallucinated" information. This paper is an experimental investigation aimed at uncovering the robustness challenges that LMs encounter when tasked with knowledge base question answering (KBQA). The investigation covers scenarios with inconsistent data distribution between training and inference, such as generalization to unseen domains, adaptation to various language variations, and transferability across different datasets. Our comprehensive experiments reveal that even when employed with our proposed data augmentation techniques, advanced small and large language models exhibit poor performance in various dimensions. While the LM is a promising technology, the robustness of the current form in dealing with complex environments is fragile and of limited practicality because of the data distribution issue. This calls for future research on data collection and LM learning paradims.
format Preprint
id arxiv_https___arxiv_org_abs_2309_08345
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Data Distribution Bottlenecks in Grounding Language Models to Knowledge Bases
Shu, Yiheng
Yu, Zhiwei
Computation and Language
Artificial Intelligence
Language models (LMs) have already demonstrated remarkable abilities in understanding and generating both natural and formal language. Despite these advances, their integration with real-world environments such as large-scale knowledge bases (KBs) remains an underdeveloped area, affecting applications such as semantic parsing and indulging in "hallucinated" information. This paper is an experimental investigation aimed at uncovering the robustness challenges that LMs encounter when tasked with knowledge base question answering (KBQA). The investigation covers scenarios with inconsistent data distribution between training and inference, such as generalization to unseen domains, adaptation to various language variations, and transferability across different datasets. Our comprehensive experiments reveal that even when employed with our proposed data augmentation techniques, advanced small and large language models exhibit poor performance in various dimensions. While the LM is a promising technology, the robustness of the current form in dealing with complex environments is fragile and of limited practicality because of the data distribution issue. This calls for future research on data collection and LM learning paradims.
title Data Distribution Bottlenecks in Grounding Language Models to Knowledge Bases
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2309.08345