CausalT5K: Diagnosing and Informing Refusal for Trustworthy Causal Reasoning of Skepticism, Sycophancy, Detection-Correction, and Rung Collapse

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Geng, Longling, Ouyang, Andy, Wu, Theodore, Barretto, Daphne, Hayes, Matthew John, Cooper, Rachael, Zeng, Yuqiao, Vijay, Sameer, Ancone, Gia, Rai, Ankit, Wolfman, Matthew, Flanagan, Patrick, Chang, Edward Y.
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911435133550592
author Geng, Longling
Ouyang, Andy
Wu, Theodore
Barretto, Daphne
Hayes, Matthew John
Cooper, Rachael
Zeng, Yuqiao
Vijay, Sameer
Ancone, Gia
Rai, Ankit
Wolfman, Matthew
Flanagan, Patrick
Chang, Edward Y.
author_facet Geng, Longling
Ouyang, Andy
Wu, Theodore
Barretto, Daphne
Hayes, Matthew John
Cooper, Rachael
Zeng, Yuqiao
Vijay, Sameer
Ancone, Gia
Rai, Ankit
Wolfman, Matthew
Flanagan, Patrick
Chang, Edward Y.
contents LLM failures in causal reasoning, including sycophancy, rung collapse, and miscalibrated refusal, are well-documented, yet progress on remediation is slow because no benchmark enables systematic diagnosis. We introduce CausalT5K, a diagnostic benchmark of over 5,000 cases across 10 domains that tests three critical capabilities: (1) detecting rung collapse, where models answer interventional queries with associational evidence; (2) resisting sycophantic drift under adversarial pressure; and (3) generating Wise Refusals that specify missing information when evidence is underdetermined. Unlike synthetic benchmarks, CausalT5K embeds causal traps in realistic narratives and decomposes performance into Utility (sensitivity) and Safety (specificity), revealing failure modes invisible to aggregate accuracy. Developed through a rigorous human-machine collaborative pipeline involving 40 domain experts, iterative cross-validation cycles, and composite verification via rule-based, LLM, and human scoring, CausalT5K implements Pearl's Ladder of Causation as research infrastructure. Preliminary experiments reveal a Four-Quadrant Control Landscape where static audit policies universally fail, a finding that demonstrates CausalT5K's value for advancing trustworthy reasoning systems. Repository: https://github.com/genglongling/CausalT5kBench
format Preprint
id arxiv_https___arxiv_org_abs_2602_08939
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle CausalT5K: Diagnosing and Informing Refusal for Trustworthy Causal Reasoning of Skepticism, Sycophancy, Detection-Correction, and Rung Collapse
Geng, Longling
Ouyang, Andy
Wu, Theodore
Barretto, Daphne
Hayes, Matthew John
Cooper, Rachael
Zeng, Yuqiao
Vijay, Sameer
Ancone, Gia
Rai, Ankit
Wolfman, Matthew
Flanagan, Patrick
Chang, Edward Y.
Artificial Intelligence
I.2.7
LLM failures in causal reasoning, including sycophancy, rung collapse, and miscalibrated refusal, are well-documented, yet progress on remediation is slow because no benchmark enables systematic diagnosis. We introduce CausalT5K, a diagnostic benchmark of over 5,000 cases across 10 domains that tests three critical capabilities: (1) detecting rung collapse, where models answer interventional queries with associational evidence; (2) resisting sycophantic drift under adversarial pressure; and (3) generating Wise Refusals that specify missing information when evidence is underdetermined. Unlike synthetic benchmarks, CausalT5K embeds causal traps in realistic narratives and decomposes performance into Utility (sensitivity) and Safety (specificity), revealing failure modes invisible to aggregate accuracy. Developed through a rigorous human-machine collaborative pipeline involving 40 domain experts, iterative cross-validation cycles, and composite verification via rule-based, LLM, and human scoring, CausalT5K implements Pearl's Ladder of Causation as research infrastructure. Preliminary experiments reveal a Four-Quadrant Control Landscape where static audit policies universally fail, a finding that demonstrates CausalT5K's value for advancing trustworthy reasoning systems. Repository: https://github.com/genglongling/CausalT5kBench
title CausalT5K: Diagnosing and Informing Refusal for Trustworthy Causal Reasoning of Skepticism, Sycophancy, Detection-Correction, and Rung Collapse
topic Artificial Intelligence
I.2.7
url https://arxiv.org/abs/2602.08939