CLadder: Assessing Causal Reasoning in Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jin, Zhijing, Chen, Yuen, Leeb, Felix, Gresele, Luigi, Kamal, Ojasv, Lyu, Zhiheng, Blin, Kevin, Adauto, Fernando Gonzalez, Kleiman-Weiner, Max, Sachan, Mrinmaya, Schölkopf, Bernhard
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913198517518336
author Jin, Zhijing
Chen, Yuen
Leeb, Felix
Gresele, Luigi
Kamal, Ojasv
Lyu, Zhiheng
Blin, Kevin
Adauto, Fernando Gonzalez
Kleiman-Weiner, Max
Sachan, Mrinmaya
Schölkopf, Bernhard
author_facet Jin, Zhijing
Chen, Yuen
Leeb, Felix
Gresele, Luigi
Kamal, Ojasv
Lyu, Zhiheng
Blin, Kevin
Adauto, Fernando Gonzalez
Kleiman-Weiner, Max
Sachan, Mrinmaya
Schölkopf, Bernhard
contents The ability to perform causal reasoning is widely considered a core feature of intelligence. In this work, we investigate whether large language models (LLMs) can coherently reason about causality. Much of the existing work in natural language processing (NLP) focuses on evaluating commonsense causal reasoning in LLMs, thus failing to assess whether a model can perform causal inference in accordance with a set of well-defined formal rules. To address this, we propose a new NLP task, causal inference in natural language, inspired by the "causal inference engine" postulated by Judea Pearl et al. We compose a large dataset, CLadder, with 10K samples: based on a collection of causal graphs and queries (associational, interventional, and counterfactual), we obtain symbolic questions and ground-truth answers, through an oracle causal inference engine. These are then translated into natural language. We evaluate multiple LLMs on our dataset, and we introduce and evaluate a bespoke chain-of-thought prompting strategy, CausalCoT. We show that our task is highly challenging for LLMs, and we conduct an in-depth analysis to gain deeper insights into the causal reasoning abilities of LLMs. Our data is open-sourced at https://huggingface.co/datasets/causalNLP/cladder, and our code can be found at https://github.com/causalNLP/cladder.
format Preprint
id arxiv_https___arxiv_org_abs_2312_04350
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle CLadder: Assessing Causal Reasoning in Language Models
Jin, Zhijing
Chen, Yuen
Leeb, Felix
Gresele, Luigi
Kamal, Ojasv
Lyu, Zhiheng
Blin, Kevin
Adauto, Fernando Gonzalez
Kleiman-Weiner, Max
Sachan, Mrinmaya
Schölkopf, Bernhard
Computation and Language
Artificial Intelligence
Machine Learning
The ability to perform causal reasoning is widely considered a core feature of intelligence. In this work, we investigate whether large language models (LLMs) can coherently reason about causality. Much of the existing work in natural language processing (NLP) focuses on evaluating commonsense causal reasoning in LLMs, thus failing to assess whether a model can perform causal inference in accordance with a set of well-defined formal rules. To address this, we propose a new NLP task, causal inference in natural language, inspired by the "causal inference engine" postulated by Judea Pearl et al. We compose a large dataset, CLadder, with 10K samples: based on a collection of causal graphs and queries (associational, interventional, and counterfactual), we obtain symbolic questions and ground-truth answers, through an oracle causal inference engine. These are then translated into natural language. We evaluate multiple LLMs on our dataset, and we introduce and evaluate a bespoke chain-of-thought prompting strategy, CausalCoT. We show that our task is highly challenging for LLMs, and we conduct an in-depth analysis to gain deeper insights into the causal reasoning abilities of LLMs. Our data is open-sourced at https://huggingface.co/datasets/causalNLP/cladder, and our code can be found at https://github.com/causalNLP/cladder.
title CLadder: Assessing Causal Reasoning in Language Models
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2312.04350