Causality can systematically address the monsters under the bench(marks)
Fuente:
arXiv
Saved in:
| Main Authors: | Leeb, Felix, Jin, Zhijing, Schölkopf, Bernhard |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Causality Is Key to Understand and Balance Multiple Goals in Trustworthy ML and Foundation Models
by: Binkyte, Ruta, et al.
Published: (2025)
by: Binkyte, Ruta, et al.
Published: (2025)
CLadder: Assessing Causal Reasoning in Language Models
by: Jin, Zhijing, et al.
Published: (2023)
by: Jin, Zhijing, et al.
Published: (2023)
Stargazer: A Scalable Model-Fitting Benchmark Environment for AI Agents under Astrophysical Constraints
by: Liu, Xinge, et al.
Published: (2026)
by: Liu, Xinge, et al.
Published: (2026)
Improving Large Language Model Safety with Contrastive Representation Learning
by: Simko, Samuel, et al.
Published: (2025)
by: Simko, Samuel, et al.
Published: (2025)
CausalCite: A Causal Formulation of Paper Citations
by: Kumar, Ishan, et al.
Published: (2023)
by: Kumar, Ishan, et al.
Published: (2023)
Causality for Natural Language Processing
by: Jin, Zhijing
Published: (2025)
by: Jin, Zhijing
Published: (2025)
Causal Responsibility Attribution for Human-AI Collaboration
by: Qi, Yahang, et al.
Published: (2024)
by: Qi, Yahang, et al.
Published: (2024)
Quriosity: Analyzing Human Questioning Behavior and Causal Inquiry through Curiosity-Driven Queries
by: Ceraolo, Roberto, et al.
Published: (2024)
by: Ceraolo, Roberto, et al.
Published: (2024)
Deep Backtracking Counterfactuals for Causally Compliant Explanations
by: Kladny, Klaus-Rudolf, et al.
Published: (2023)
by: Kladny, Klaus-Rudolf, et al.
Published: (2023)
Identifiable Exchangeable Mechanisms for Causal Structure and Representation Learning
by: Reizinger, Patrik, et al.
Published: (2024)
by: Reizinger, Patrik, et al.
Published: (2024)
Learning Nonlinear Causal Reductions to Explain Reinforcement Learning Policies
by: Kekić, Armin, et al.
Published: (2025)
by: Kekić, Armin, et al.
Published: (2025)
Causal vs. Anticausal merging of predictors
by: Mejia, Sergio Hernan Garrido, et al.
Published: (2025)
by: Mejia, Sergio Hernan Garrido, et al.
Published: (2025)
Causal Component Analysis
by: Wendong, Liang, et al.
Published: (2023)
by: Wendong, Liang, et al.
Published: (2023)
Can Large Language Models Infer Causation from Correlation?
by: Jin, Zhijing, et al.
Published: (2023)
by: Jin, Zhijing, et al.
Published: (2023)
Analyzing the Role of Semantic Representations in the Era of Large Language Models
by: Jin, Zhijing, et al.
Published: (2024)
by: Jin, Zhijing, et al.
Published: (2024)
A Probabilistic Model Behind Self-Supervised Learning
by: Bizeul, Alice, et al.
Published: (2024)
by: Bizeul, Alice, et al.
Published: (2024)
Out-of-Variable Generalization for Discriminative Models
by: Guo, Siyuan, et al.
Published: (2023)
by: Guo, Siyuan, et al.
Published: (2023)
Computational Arbitrage in AI Model Markets
by: Olmedo, Ricardo, et al.
Published: (2026)
by: Olmedo, Ricardo, et al.
Published: (2026)
Learning Interpretable Concepts: Unifying Causal Representation Learning and Foundation Models
by: Rajendran, Goutham, et al.
Published: (2024)
by: Rajendran, Goutham, et al.
Published: (2024)
Conformal Generative Modeling with Improved Sample Efficiency through Sequential Greedy Filtering
by: Kladny, Klaus-Rudolf, et al.
Published: (2024)
by: Kladny, Klaus-Rudolf, et al.
Published: (2024)
Riemannian Networks over Full-Rank Correlation Matrices
by: Chen, Ziheng, et al.
Published: (2026)
by: Chen, Ziheng, et al.
Published: (2026)
Preference Elicitation for Offline Reinforcement Learning
by: Pace, Alizée, et al.
Published: (2024)
by: Pace, Alizée, et al.
Published: (2024)
Provable Privacy with Non-Private Pre-Processing
by: Hu, Yaxi, et al.
Published: (2024)
by: Hu, Yaxi, et al.
Published: (2024)
Identifying Intervenable and Interpretable Features via Orthogonality Regularization
by: Miller, Moritz, et al.
Published: (2026)
by: Miller, Moritz, et al.
Published: (2026)
Riemannian Batch Normalization: A Gyro Approach
by: Chen, Ziheng, et al.
Published: (2025)
by: Chen, Ziheng, et al.
Published: (2025)
Trustworthy AI Suffers from Invariance Conflicts and Causality is The Solution
by: Binkyte, Ruta, et al.
Published: (2026)
by: Binkyte, Ruta, et al.
Published: (2026)
Natural Building Blocks for Structured World Models: Theory, Evidence, and Scaling
by: Da Costa, Lancelot, et al.
Published: (2025)
by: Da Costa, Lancelot, et al.
Published: (2025)
CORE: Measuring Multi-Agent LLM Interaction Quality under Game-Theoretic Pressures
by: Pandey, Punya Syon, et al.
Published: (2025)
by: Pandey, Punya Syon, et al.
Published: (2025)
Intrinsically Interpretable Attention via Sparse Post-Training
by: Draye, Florent, et al.
Published: (2025)
by: Draye, Florent, et al.
Published: (2025)
Counterfactual reasoning: an analysis of in-context emergence
by: Miller, Moritz, et al.
Published: (2025)
by: Miller, Moritz, et al.
Published: (2025)
Algorithmic causal structure emerging through compression
by: Wendong, Liang, et al.
Published: (2025)
by: Wendong, Liang, et al.
Published: (2025)
Hyperbolic Busemann Neural Networks
by: Chen, Ziheng, et al.
Published: (2026)
by: Chen, Ziheng, et al.
Published: (2026)
BinaryPPO: Efficient Policy Optimization for Binary Classification
by: Pandey, Punya Syon, et al.
Published: (2026)
by: Pandey, Punya Syon, et al.
Published: (2026)
Implicit Personalization in Language Models: A Systematic Study
by: Jin, Zhijing, et al.
Published: (2024)
by: Jin, Zhijing, et al.
Published: (2024)
Skill Learning via Policy Diversity Yields Identifiable Representations for Reinforcement Learning
by: Reizinger, Patrik, et al.
Published: (2025)
by: Reizinger, Patrik, et al.
Published: (2025)
Accuracy on the wrong line: On the pitfalls of noisy data for out-of-distribution generalisation
by: Sanyal, Amartya, et al.
Published: (2024)
by: Sanyal, Amartya, et al.
Published: (2024)
AI Research Agents for Machine Learning: Search, Exploration, and Generalization in MLE-bench
by: Toledo, Edan, et al.
Published: (2025)
by: Toledo, Edan, et al.
Published: (2025)
Difficulty-Based Preference Data Selection by DPO Implicit Reward Gap
by: Qi, Xuan, et al.
Published: (2025)
by: Qi, Xuan, et al.
Published: (2025)
Adaptable Cardiovascular Disease Risk Prediction from Heterogeneous Data using Large Language Models
by: Lübeck, Frederike, et al.
Published: (2025)
by: Lübeck, Frederike, et al.
Published: (2025)
MathGAP: Out-of-Distribution Evaluation on Problems with Arbitrarily Complex Proofs
by: Opedal, Andreas, et al.
Published: (2024)
by: Opedal, Andreas, et al.
Published: (2024)
Similar Items
-
Causality Is Key to Understand and Balance Multiple Goals in Trustworthy ML and Foundation Models
by: Binkyte, Ruta, et al.
Published: (2025) -
CLadder: Assessing Causal Reasoning in Language Models
by: Jin, Zhijing, et al.
Published: (2023) -
Stargazer: A Scalable Model-Fitting Benchmark Environment for AI Agents under Astrophysical Constraints
by: Liu, Xinge, et al.
Published: (2026) -
Improving Large Language Model Safety with Contrastive Representation Learning
by: Simko, Samuel, et al.
Published: (2025) -
CausalCite: A Causal Formulation of Paper Citations
by: Kumar, Ishan, et al.
Published: (2023)