CKBP v2: Better Annotation and Reasoning for Commonsense Knowledge Base Population

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fang, Tianqing, Do, Quyet V., Zheng, Zihao, Wang, Weiqi, Choi, Sehyun, Wang, Zhaowei, Song, Yangqiu
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910615365222400
author Fang, Tianqing
Do, Quyet V.
Zheng, Zihao
Wang, Weiqi
Choi, Sehyun
Wang, Zhaowei
Song, Yangqiu
author_facet Fang, Tianqing
Do, Quyet V.
Zheng, Zihao
Wang, Weiqi
Choi, Sehyun
Wang, Zhaowei
Song, Yangqiu
contents Commonsense Knowledge Bases (CSKB) Population, which aims at automatically expanding knowledge in CSKBs with external resources, is an important yet hard task in NLP. Fang et al. (2021a) proposed a CSKB Population (CKBP) framework with an evaluation set CKBP v1. However, CKBP v1 relies on crowdsourced annotations that suffer from a considerable number of mislabeled answers, and the evaluationset lacks alignment with the external knowledge source due to random sampling. In this paper, we introduce CKBP v2, a new high-quality CSKB Population evaluation set that addresses the two aforementioned issues by employing domain experts as annotators and incorporating diversified adversarial samples to make the evaluation data more representative. We show that CKBP v2 serves as a challenging and representative evaluation dataset for the CSKB Population task, while its development set aids in selecting a population model that leads to improved knowledge acquisition for downstream commonsense reasoning. A better population model can also help acquire more informative commonsense knowledge as additional supervision signals for both generative commonsense inference and zero-shot commonsense question answering. Specifically, the question-answering model based on DeBERTa-v3-large (He et al., 2023b) even outperforms powerful large language models in a zero-shot setting, including ChatGPT and GPT-3.5.
format Preprint
id arxiv_https___arxiv_org_abs_2304_10392
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle CKBP v2: Better Annotation and Reasoning for Commonsense Knowledge Base Population
Fang, Tianqing
Do, Quyet V.
Zheng, Zihao
Wang, Weiqi
Choi, Sehyun
Wang, Zhaowei
Song, Yangqiu
Computation and Language
Commonsense Knowledge Bases (CSKB) Population, which aims at automatically expanding knowledge in CSKBs with external resources, is an important yet hard task in NLP. Fang et al. (2021a) proposed a CSKB Population (CKBP) framework with an evaluation set CKBP v1. However, CKBP v1 relies on crowdsourced annotations that suffer from a considerable number of mislabeled answers, and the evaluationset lacks alignment with the external knowledge source due to random sampling. In this paper, we introduce CKBP v2, a new high-quality CSKB Population evaluation set that addresses the two aforementioned issues by employing domain experts as annotators and incorporating diversified adversarial samples to make the evaluation data more representative. We show that CKBP v2 serves as a challenging and representative evaluation dataset for the CSKB Population task, while its development set aids in selecting a population model that leads to improved knowledge acquisition for downstream commonsense reasoning. A better population model can also help acquire more informative commonsense knowledge as additional supervision signals for both generative commonsense inference and zero-shot commonsense question answering. Specifically, the question-answering model based on DeBERTa-v3-large (He et al., 2023b) even outperforms powerful large language models in a zero-shot setting, including ChatGPT and GPT-3.5.
title CKBP v2: Better Annotation and Reasoning for Commonsense Knowledge Base Population
topic Computation and Language
url https://arxiv.org/abs/2304.10392