A Critical Review of Causal Reasoning Benchmarks for Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Linying, Shirvaikar, Vik, Clivio, Oscar, Falck, Fabian
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929417516744704
author Yang, Linying
Shirvaikar, Vik
Clivio, Oscar
Falck, Fabian
author_facet Yang, Linying
Shirvaikar, Vik
Clivio, Oscar
Falck, Fabian
contents Numerous benchmarks aim to evaluate the capabilities of Large Language Models (LLMs) for causal inference and reasoning. However, many of them can likely be solved through the retrieval of domain knowledge, questioning whether they achieve their purpose. In this review, we present a comprehensive overview of LLM benchmarks for causality. We highlight how recent benchmarks move towards a more thorough definition of causal reasoning by incorporating interventional or counterfactual reasoning. We derive a set of criteria that a useful benchmark or set of benchmarks should aim to satisfy. We hope this work will pave the way towards a general framework for the assessment of causal understanding in LLMs and the design of novel benchmarks.
format Preprint
id arxiv_https___arxiv_org_abs_2407_08029
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle A Critical Review of Causal Reasoning Benchmarks for Large Language Models
Yang, Linying
Shirvaikar, Vik
Clivio, Oscar
Falck, Fabian
Machine Learning
Computation and Language
Numerous benchmarks aim to evaluate the capabilities of Large Language Models (LLMs) for causal inference and reasoning. However, many of them can likely be solved through the retrieval of domain knowledge, questioning whether they achieve their purpose. In this review, we present a comprehensive overview of LLM benchmarks for causality. We highlight how recent benchmarks move towards a more thorough definition of causal reasoning by incorporating interventional or counterfactual reasoning. We derive a set of criteria that a useful benchmark or set of benchmarks should aim to satisfy. We hope this work will pave the way towards a general framework for the assessment of causal understanding in LLMs and the design of novel benchmarks.
title A Critical Review of Causal Reasoning Benchmarks for Large Language Models
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2407.08029