Direct Evaluation of Chain-of-Thought in Multi-hop Reasoning with Knowledge Graphs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Nguyen, Minh-Vuong, Luo, Linhao, Shiri, Fatemeh, Phung, Dinh, Li, Yuan-Fang, Vu, Thuy-Trang, Haffari, Gholamreza
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929391648374784
author Nguyen, Minh-Vuong
Luo, Linhao
Shiri, Fatemeh
Phung, Dinh
Li, Yuan-Fang
Vu, Thuy-Trang
Haffari, Gholamreza
author_facet Nguyen, Minh-Vuong
Luo, Linhao
Shiri, Fatemeh
Phung, Dinh
Li, Yuan-Fang
Vu, Thuy-Trang
Haffari, Gholamreza
contents Large language models (LLMs) demonstrate strong reasoning abilities when prompted to generate chain-of-thought (CoT) explanations alongside answers. However, previous research on evaluating LLMs has solely focused on answer accuracy, neglecting the correctness of the generated CoT. In this paper, we delve deeper into the CoT reasoning capabilities of LLMs in multi-hop question answering by utilizing knowledge graphs (KGs). We propose a novel discriminative and generative CoT evaluation paradigm to assess LLMs' knowledge of reasoning and the accuracy of the generated CoT. Through experiments conducted on 5 different families of LLMs across 2 multi-hop question-answering datasets, we find that LLMs possess sufficient knowledge to perform reasoning. However, there exists a significant disparity between answer accuracy and faithfulness of the CoT reasoning generated by LLMs, indicating that they often arrive at correct answers through incorrect reasoning.
format Preprint
id arxiv_https___arxiv_org_abs_2402_11199
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Direct Evaluation of Chain-of-Thought in Multi-hop Reasoning with Knowledge Graphs
Nguyen, Minh-Vuong
Luo, Linhao
Shiri, Fatemeh
Phung, Dinh
Li, Yuan-Fang
Vu, Thuy-Trang
Haffari, Gholamreza
Computation and Language
Large language models (LLMs) demonstrate strong reasoning abilities when prompted to generate chain-of-thought (CoT) explanations alongside answers. However, previous research on evaluating LLMs has solely focused on answer accuracy, neglecting the correctness of the generated CoT. In this paper, we delve deeper into the CoT reasoning capabilities of LLMs in multi-hop question answering by utilizing knowledge graphs (KGs). We propose a novel discriminative and generative CoT evaluation paradigm to assess LLMs' knowledge of reasoning and the accuracy of the generated CoT. Through experiments conducted on 5 different families of LLMs across 2 multi-hop question-answering datasets, we find that LLMs possess sufficient knowledge to perform reasoning. However, there exists a significant disparity between answer accuracy and faithfulness of the CoT reasoning generated by LLMs, indicating that they often arrive at correct answers through incorrect reasoning.
title Direct Evaluation of Chain-of-Thought in Multi-hop Reasoning with Knowledge Graphs
topic Computation and Language
url https://arxiv.org/abs/2402.11199