Nuance Matters: Probing Epistemic Consistency in Causal Reasoning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cui, Shaobo, Li, Junyou, Mouchel, Luca, Feng, Yiyang, Faltings, Boi
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912009917825024
author Cui, Shaobo
Li, Junyou
Mouchel, Luca
Feng, Yiyang
Faltings, Boi
author_facet Cui, Shaobo
Li, Junyou
Mouchel, Luca
Feng, Yiyang
Faltings, Boi
contents To address this gap, our study introduces the concept of causal epistemic consistency, which focuses on the self-consistency of Large Language Models (LLMs) in differentiating intermediates with nuanced differences in causal reasoning. We propose a suite of novel metrics -- intensity ranking concordance, cross-group position agreement, and intra-group clustering -- to evaluate LLMs on this front. Through extensive empirical studies on 21 high-profile LLMs, including GPT-4, Claude3, and LLaMA3-70B, we have favoring evidence that current models struggle to maintain epistemic consistency in identifying the polarity and intensity of intermediates in causal reasoning. Additionally, we explore the potential of using internal token probabilities as an auxiliary tool to maintain causal epistemic consistency. In summary, our study bridges a critical gap in AI research by investigating the self-consistency over fine-grained intermediates involved in causal reasoning.
format Preprint
id arxiv_https___arxiv_org_abs_2409_00103
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Nuance Matters: Probing Epistemic Consistency in Causal Reasoning
Cui, Shaobo
Li, Junyou
Mouchel, Luca
Feng, Yiyang
Faltings, Boi
Computation and Language
Artificial Intelligence
To address this gap, our study introduces the concept of causal epistemic consistency, which focuses on the self-consistency of Large Language Models (LLMs) in differentiating intermediates with nuanced differences in causal reasoning. We propose a suite of novel metrics -- intensity ranking concordance, cross-group position agreement, and intra-group clustering -- to evaluate LLMs on this front. Through extensive empirical studies on 21 high-profile LLMs, including GPT-4, Claude3, and LLaMA3-70B, we have favoring evidence that current models struggle to maintain epistemic consistency in identifying the polarity and intensity of intermediates in causal reasoning. Additionally, we explore the potential of using internal token probabilities as an auxiliary tool to maintain causal epistemic consistency. In summary, our study bridges a critical gap in AI research by investigating the self-consistency over fine-grained intermediates involved in causal reasoning.
title Nuance Matters: Probing Epistemic Consistency in Causal Reasoning
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2409.00103