METER: Evaluating Multi-Level Contextual Causal Reasoning in Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Pengfeng, Huang, Chen, Hao, Chaoqun, Chen, Hongyao, Wei, Xiao-Yong, Lei, Wenqiang, Ng, See-Kiong
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908969440641024
author Li, Pengfeng
Huang, Chen
Hao, Chaoqun
Chen, Hongyao
Wei, Xiao-Yong
Lei, Wenqiang
Ng, See-Kiong
author_facet Li, Pengfeng
Huang, Chen
Hao, Chaoqun
Chen, Hongyao
Wei, Xiao-Yong
Lei, Wenqiang
Ng, See-Kiong
contents Contextual causal reasoning is a critical yet challenging capability for Large Language Models (LLMs). Existing benchmarks, however, often evaluate this skill in fragmented settings, failing to ensure context consistency or cover the full causal hierarchy. To address this, we pioneer METER to systematically benchmark LLMs across all three levels of the causal ladder under a unified context setting. Our extensive evaluation of various LLMs reveals a significant decline in proficiency as tasks ascend the causal hierarchy. To diagnose this degradation, we conduct a deep mechanistic analysis via both error pattern identification and internal information flow tracing. Our analysis reveals two primary failure modes: (1) LLMs are susceptible to distraction by causally irrelevant but factually correct information at lower level of causality; and (2) as tasks ascend the causal hierarchy, faithfulness to the provided context degrades, leading to a reduced performance. We belive our work advances our understanding of the mechanisms behind LLM contextual causal reasoning and establishes a critical foundation for future research. Our code and dataset are available at https://github.com/SCUNLP/METER .
format Preprint
id arxiv_https___arxiv_org_abs_2604_11502
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle METER: Evaluating Multi-Level Contextual Causal Reasoning in Large Language Models
Li, Pengfeng
Huang, Chen
Hao, Chaoqun
Chen, Hongyao
Wei, Xiao-Yong
Lei, Wenqiang
Ng, See-Kiong
Computation and Language
Artificial Intelligence
Contextual causal reasoning is a critical yet challenging capability for Large Language Models (LLMs). Existing benchmarks, however, often evaluate this skill in fragmented settings, failing to ensure context consistency or cover the full causal hierarchy. To address this, we pioneer METER to systematically benchmark LLMs across all three levels of the causal ladder under a unified context setting. Our extensive evaluation of various LLMs reveals a significant decline in proficiency as tasks ascend the causal hierarchy. To diagnose this degradation, we conduct a deep mechanistic analysis via both error pattern identification and internal information flow tracing. Our analysis reveals two primary failure modes: (1) LLMs are susceptible to distraction by causally irrelevant but factually correct information at lower level of causality; and (2) as tasks ascend the causal hierarchy, faithfulness to the provided context degrades, leading to a reduced performance. We belive our work advances our understanding of the mechanisms behind LLM contextual causal reasoning and establishes a critical foundation for future research. Our code and dataset are available at https://github.com/SCUNLP/METER .
title METER: Evaluating Multi-Level Contextual Causal Reasoning in Large Language Models
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2604.11502