A Comprehensive Evaluation on Event Reasoning of Large Language Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Tao, Zhengwei, Jin, Zhi, Zhang, Yifan, Chen, Xiancai, Zhao, Haiyan, Li, Jia, Liang, Bing, Tao, Chongyang, Liu, Qun, Wong, Kam-Fai
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866929446289670144
author Tao, Zhengwei
Jin, Zhi
Zhang, Yifan
Chen, Xiancai
Zhao, Haiyan
Li, Jia
Liang, Bing
Tao, Chongyang
Liu, Qun
Wong, Kam-Fai
author_facet Tao, Zhengwei
Jin, Zhi
Zhang, Yifan
Chen, Xiancai
Zhao, Haiyan
Li, Jia
Liang, Bing
Tao, Chongyang
Liu, Qun
Wong, Kam-Fai
contents Event reasoning is a fundamental ability that underlies many applications. It requires event schema knowledge to perform global reasoning and needs to deal with the diversity of the inter-event relations and the reasoning paradigms. How well LLMs accomplish event reasoning on various relations and reasoning paradigms remains unknown. To mitigate this disparity, we comprehensively evaluate the abilities of event reasoning of LLMs. We introduce a novel benchmark EV2 for EValuation of EVent reasoning. EV2 consists of two levels of evaluation of schema and instance and is comprehensive in relations and reasoning paradigms. We conduct extensive experiments on EV2. We find that LLMs have abilities to accomplish event reasoning but their performances are far from satisfactory. We also notice the imbalance of event reasoning abilities in LLMs. Besides, LLMs have event schema knowledge, however, they're not aligned with humans on how to utilize the knowledge. Based on these findings, we guide the LLMs in utilizing the event schema knowledge as memory leading to improvements on event reasoning.
format Preprint
id arxiv_https___arxiv_org_abs_2404_17513
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle A Comprehensive Evaluation on Event Reasoning of Large Language Models
Tao, Zhengwei
Jin, Zhi
Zhang, Yifan
Chen, Xiancai
Zhao, Haiyan
Li, Jia
Liang, Bing
Tao, Chongyang
Liu, Qun
Wong, Kam-Fai
Computation and Language
Artificial Intelligence
Event reasoning is a fundamental ability that underlies many applications. It requires event schema knowledge to perform global reasoning and needs to deal with the diversity of the inter-event relations and the reasoning paradigms. How well LLMs accomplish event reasoning on various relations and reasoning paradigms remains unknown. To mitigate this disparity, we comprehensively evaluate the abilities of event reasoning of LLMs. We introduce a novel benchmark EV2 for EValuation of EVent reasoning. EV2 consists of two levels of evaluation of schema and instance and is comprehensive in relations and reasoning paradigms. We conduct extensive experiments on EV2. We find that LLMs have abilities to accomplish event reasoning but their performances are far from satisfactory. We also notice the imbalance of event reasoning abilities in LLMs. Besides, LLMs have event schema knowledge, however, they're not aligned with humans on how to utilize the knowledge. Based on these findings, we guide the LLMs in utilizing the event schema knowledge as memory leading to improvements on event reasoning.
title A Comprehensive Evaluation on Event Reasoning of Large Language Models
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2404.17513