CausalReasoningBenchmark: A Real-World Benchmark for Disentangled Evaluation of Causal Identification and Estimation
Fuente:
arXiv
Guardado en:
| Autores principales: | Sawarni, Ayush, Tan, Jiyuan, Syrgkanis, Vasilis |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Consistency of Neural Causal Partial Identification
por: Tan, Jiyuan, et al.
Publicado: (2024)
por: Tan, Jiyuan, et al.
Publicado: (2024)
Preference Learning with Response Time: Robust Losses and Guarantees
por: Sawarni, Ayush, et al.
Publicado: (2025)
por: Sawarni, Ayush, et al.
Publicado: (2025)
Learning Causal Representations from General Environments: Identifiability and Intrinsic Ambiguity
por: Jin, Jikai, et al.
Publicado: (2023)
por: Jin, Jikai, et al.
Publicado: (2023)
Empirical Analysis of Model Selection for Heterogeneous Causal Effect Estimation
por: Mahajan, Divyat, et al.
Publicado: (2022)
por: Mahajan, Divyat, et al.
Publicado: (2022)
Statistical Inference and Learning for Shapley Additive Explanations (SHAP)
por: Whitehouse, Justin, et al.
Publicado: (2026)
por: Whitehouse, Justin, et al.
Publicado: (2026)
Discovering Hierarchical Latent Capabilities of Language Models via Causal Representation Learning
por: Jin, Jikai, et al.
Publicado: (2025)
por: Jin, Jikai, et al.
Publicado: (2025)
Adaptive Estimation and Inference in Conditional Moment Models via the Discrepancy Principle
por: Tan, Jiyuan, et al.
Publicado: (2026)
por: Tan, Jiyuan, et al.
Publicado: (2026)
Policy Learning with Abstention
por: Sawarni, Ayush, et al.
Publicado: (2025)
por: Sawarni, Ayush, et al.
Publicado: (2025)
NoisyCausal: A Benchmark for Evaluating Causal Reasoning Under Structured Noise
por: Xu, Zhi, et al.
Publicado: (2026)
por: Xu, Zhi, et al.
Publicado: (2026)
CausalStep: A Benchmark for Explicit Stepwise Causal Reasoning in Videos
por: Li, Xuchen, et al.
Publicado: (2025)
por: Li, Xuchen, et al.
Publicado: (2025)
ACCESS : A Benchmark for Abstract Causal Event Discovery and Reasoning
por: Vo, Vy, et al.
Publicado: (2025)
por: Vo, Vy, et al.
Publicado: (2025)
Partial Identification of Policy-Relevant Treatment Effects with Instrumental Variables via Optimal Transport
por: Tan, Jiyuan, et al.
Publicado: (2026)
por: Tan, Jiyuan, et al.
Publicado: (2026)
Towards a Benchmark for Causal Business Process Reasoning with LLMs
por: Fournier, Fabiana, et al.
Publicado: (2024)
por: Fournier, Fabiana, et al.
Publicado: (2024)
Estimation of Treatment Effects in Extreme and Unobserved Data
por: Tan, Jiyuan, et al.
Publicado: (2025)
por: Tan, Jiyuan, et al.
Publicado: (2025)
Envision: Benchmarking Unified Understanding & Generation for Causal World Process Insights
por: Tian, Juanxi, et al.
Publicado: (2025)
por: Tian, Juanxi, et al.
Publicado: (2025)
CausalVLBench: Benchmarking Visual Causal Reasoning in Large Vision-Language Models
por: Komanduri, Aneesh, et al.
Publicado: (2025)
por: Komanduri, Aneesh, et al.
Publicado: (2025)
InterveneBench: Benchmarking LLMs for Intervention Reasoning and Causal Study Design in Real Social Systems
por: Shi, Shaojie, et al.
Publicado: (2026)
por: Shi, Shaojie, et al.
Publicado: (2026)
Direct Preference Optimization with Unobserved Preference Heterogeneity: The Necessity of Ternary Preferences
por: Chidambaram, Keertana, et al.
Publicado: (2025)
por: Chidambaram, Keertana, et al.
Publicado: (2025)
CausalProfiler: Generating Synthetic Benchmarks for Rigorous and Transparent Evaluation of Causal Machine Learning
por: Panayiotou, Panayiotis, et al.
Publicado: (2025)
por: Panayiotou, Panayiotis, et al.
Publicado: (2025)
CausalFlip: A Benchmark for LLM Causal Judgment Beyond Semantic Matching
por: Wang, Yuzhe, et al.
Publicado: (2026)
por: Wang, Yuzhe, et al.
Publicado: (2026)
OCDB: Revisiting Causal Discovery with a Comprehensive Benchmark and Evaluation Framework
por: Zhou, Wei, et al.
Publicado: (2024)
por: Zhou, Wei, et al.
Publicado: (2024)
Causal Q-Aggregation for CATE Model Selection
por: Lan, Hui, et al.
Publicado: (2023)
por: Lan, Hui, et al.
Publicado: (2023)
CausalVQA: A Physically Grounded Causal Reasoning Benchmark for Video Models
por: Foss, Aaron, et al.
Publicado: (2025)
por: Foss, Aaron, et al.
Publicado: (2025)
CausalARC: Abstract Reasoning with Causal World Models
por: Maasch, Jacqueline, et al.
Publicado: (2025)
por: Maasch, Jacqueline, et al.
Publicado: (2025)
TIME: A Multi-level Benchmark for Temporal Reasoning of LLMs in Real-World Scenarios
por: Wei, Shaohang, et al.
Publicado: (2025)
por: Wei, Shaohang, et al.
Publicado: (2025)
Bucketing the Good Apples: A Method for Diagnosing and Improving Causal Abstraction
por: Puyin, Li, et al.
Publicado: (2026)
por: Puyin, Li, et al.
Publicado: (2026)
Disentangled Double Machine Learning for Accurate Causal Effect Estimation
por: Xiang, Guodu, et al.
Publicado: (2026)
por: Xiang, Guodu, et al.
Publicado: (2026)
Hybrid Causal Identification and Causal Mechanism Clustering
por: Liu, Saixiong, et al.
Publicado: (2025)
por: Liu, Saixiong, et al.
Publicado: (2025)
Reframing Spatial Reasoning Evaluation in Language Models: A Real-World Simulation Benchmark for Qualitative Reasoning
por: Li, Fangjun, et al.
Publicado: (2024)
por: Li, Fangjun, et al.
Publicado: (2024)
EconCausal: A Context-Aware Economic Reasoning Benchmark for Large Language Models
por: Lee, Donggyu, et al.
Publicado: (2025)
por: Lee, Donggyu, et al.
Publicado: (2025)
Are LLMs Capable of Data-based Statistical and Causal Reasoning? Benchmarking Advanced Quantitative Reasoning with Data
por: Liu, Xiao, et al.
Publicado: (2024)
por: Liu, Xiao, et al.
Publicado: (2024)
Causality for Tabular Data Synthesis: A High-Order Structure Causal Benchmark Framework
por: Tu, Ruibo, et al.
Publicado: (2024)
por: Tu, Ruibo, et al.
Publicado: (2024)
Multimodal Causal Reasoning Benchmark: Challenging Vision Large Language Models to Discern Causal Links Across Modalities
por: Li, Zhiyuan, et al.
Publicado: (2024)
por: Li, Zhiyuan, et al.
Publicado: (2024)
Towards efficient representation identification in supervised learning
por: Ahuja, Kartik, et al.
Publicado: (2022)
por: Ahuja, Kartik, et al.
Publicado: (2022)
Causal-LLaVA: Causal Disentanglement for Mitigating Hallucination in Multimodal Large Language Models
por: Hu, Xinmiao, et al.
Publicado: (2025)
por: Hu, Xinmiao, et al.
Publicado: (2025)
Experimental Evaluation of ROS-Causal in Real-World Human-Robot Spatial Interaction Scenarios
por: Castri, Luca, et al.
Publicado: (2024)
por: Castri, Luca, et al.
Publicado: (2024)
Disentangled Representations for Causal Cognition
por: Torresan, Filippo, et al.
Publicado: (2024)
por: Torresan, Filippo, et al.
Publicado: (2024)
C2-Faith: Benchmarking LLM Judges for Causal and Coverage Faithfulness in Chain-of-Thought Reasoning
por: Mittal, Avni, et al.
Publicado: (2026)
por: Mittal, Avni, et al.
Publicado: (2026)
A Benchmark of Causal vs. Correlation AI for Predictive Maintenance
por: Dhande, Shaunak, et al.
Publicado: (2025)
por: Dhande, Shaunak, et al.
Publicado: (2025)
Com$^2$: A Causal-Guided Benchmark for Exploring Complex Commonsense Reasoning in Large Language Models
por: Xiong, Kai, et al.
Publicado: (2025)
por: Xiong, Kai, et al.
Publicado: (2025)
Ejemplares similares
-
Consistency of Neural Causal Partial Identification
por: Tan, Jiyuan, et al.
Publicado: (2024) -
Preference Learning with Response Time: Robust Losses and Guarantees
por: Sawarni, Ayush, et al.
Publicado: (2025) -
Learning Causal Representations from General Environments: Identifiability and Intrinsic Ambiguity
por: Jin, Jikai, et al.
Publicado: (2023) -
Empirical Analysis of Model Selection for Heterogeneous Causal Effect Estimation
por: Mahajan, Divyat, et al.
Publicado: (2022) -
Statistical Inference and Learning for Shapley Additive Explanations (SHAP)
por: Whitehouse, Justin, et al.
Publicado: (2026)