Modeling Discrimination with Causal Abstraction
Fuente:
arXiv
Saved in:
| Main Authors: | Mossé, Milan, Schechtman, Kara, Eberhardt, Frederick, Icard, Thomas |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
On Probabilistic and Causal Reasoning with Summation Operators
by: Ibeling, Duligur, et al.
Published: (2024)
by: Ibeling, Duligur, et al.
Published: (2024)
How Causal Abstraction Underpins Computational Explanation
by: Geiger, Atticus, et al.
Published: (2025)
by: Geiger, Atticus, et al.
Published: (2025)
Bucketing the Good Apples: A Method for Diagnosing and Improving Causal Abstraction
by: Puyin, Li, et al.
Published: (2026)
by: Puyin, Li, et al.
Published: (2026)
Belief in the Machine: Investigating Epistemological Blind Spots of Language Models
by: Suzgun, Mirac, et al.
Published: (2024)
by: Suzgun, Mirac, et al.
Published: (2024)
QueerBench: Quantifying Discrimination in Language Models Toward Queer Identities
by: Sosto, Mae, et al.
Published: (2024)
by: Sosto, Mae, et al.
Published: (2024)
Towards Effective Discrimination Testing for Generative AI
by: Zollo, Thomas P., et al.
Published: (2024)
by: Zollo, Thomas P., et al.
Published: (2024)
Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability
by: Geiger, Atticus, et al.
Published: (2023)
by: Geiger, Atticus, et al.
Published: (2023)
Unlawful Proxy Discrimination: A Framework for Challenging Inherently Discriminatory Algorithms
by: Weerts, Hilde, et al.
Published: (2024)
by: Weerts, Hilde, et al.
Published: (2024)
Internal Causal Mechanisms Robustly Predict Language Model Out-of-Distribution Behaviors
by: Huang, Jing, et al.
Published: (2025)
by: Huang, Jing, et al.
Published: (2025)
Societal Capacity Assessment Framework: Measuring Resilience to Inform Advanced AI Risk Management
by: Gandhi, Milan, et al.
Published: (2025)
by: Gandhi, Milan, et al.
Published: (2025)
Generative Discrimination: What Happens When Generative AI Exhibits Bias, and What Can Be Done About It
by: Hacker, Philipp
Published: (2024)
by: Hacker, Philipp
Published: (2024)
The Impact of Item-Writing Flaws on Difficulty and Discrimination in Item Response Theory
by: Schmucker, Robin, et al.
Published: (2025)
by: Schmucker, Robin, et al.
Published: (2025)
The Dual Impact of Virtual Reality: Examining the Addictive Potential and Therapeutic Applications of Immersive Media in the Metaverse
by: Bojic, Ljubisa, et al.
Published: (2024)
by: Bojic, Ljubisa, et al.
Published: (2024)
Does GPT-4 surpass human performance in linguistic pragmatics?
by: Bojic, Ljubisa, et al.
Published: (2023)
by: Bojic, Ljubisa, et al.
Published: (2023)
Causality in the Can: Diet Coke's Impact on Fatness
by: Qi, Yicheng, et al.
Published: (2024)
by: Qi, Yicheng, et al.
Published: (2024)
Toward A Causal Framework for Modeling Perception
by: Alvarez, Jose M., et al.
Published: (2024)
by: Alvarez, Jose M., et al.
Published: (2024)
Human Attribution of Causality to AI Across Agency, Misuse, and Misalignment
by: Carro, Maria Victoria, et al.
Published: (2026)
by: Carro, Maria Victoria, et al.
Published: (2026)
ASCenD-BDS: Adaptable, Stochastic and Context-aware framework for Detection of Bias, Discrimination and Stereotyping
by: Bahl, Rajiv, et al.
Published: (2025)
by: Bahl, Rajiv, et al.
Published: (2025)
LLM-Driven Robots Risk Enacting Discrimination, Violence, and Unlawful Actions
by: Hundt, Andrew, et al.
Published: (2024)
by: Hundt, Andrew, et al.
Published: (2024)
Persona-Based Simulation of Human Opinion at Population Scale
by: Li, Mao, et al.
Published: (2026)
by: Li, Mao, et al.
Published: (2026)
Evaluating Large Language Models Against Human Annotators in Latent Content Analysis: Sentiment, Political Leaning, Emotional Intensity, and Sarcasm
by: Bojic, Ljubisa, et al.
Published: (2025)
by: Bojic, Ljubisa, et al.
Published: (2025)
Who Would Chatbots Vote For? Political Preferences of ChatGPT and Gemini in the 2024 European Union Elections
by: Haman, Michael, et al.
Published: (2024)
by: Haman, Michael, et al.
Published: (2024)
PopResume: Causal Fairness Evaluation of LLM/VLM Resume Screeners with Population-Representative Dataset
by: Yu, Sumin, et al.
Published: (2026)
by: Yu, Sumin, et al.
Published: (2026)
InterveneBench: Benchmarking LLMs for Intervention Reasoning and Causal Study Design in Real Social Systems
by: Shi, Shaojie, et al.
Published: (2026)
by: Shi, Shaojie, et al.
Published: (2026)
Causal machine learning for sustainable agroecosystems
by: Sitokonstantinou, Vasileios, et al.
Published: (2024)
by: Sitokonstantinou, Vasileios, et al.
Published: (2024)
Responsible Evaluation of AI for Mental Health
by: Arnaout, Hiba, et al.
Published: (2026)
by: Arnaout, Hiba, et al.
Published: (2026)
Comprehensive AI governance requires addressing non-model gains
by: Goemans, Arthur, et al.
Published: (2026)
by: Goemans, Arthur, et al.
Published: (2026)
Tuning Derivatives for Causal Fairness in Machine Learning
by: Edström, Filip, et al.
Published: (2026)
by: Edström, Filip, et al.
Published: (2026)
Causality for Natural Language Processing
by: Jin, Zhijing
Published: (2025)
by: Jin, Zhijing
Published: (2025)
AI reasoning effort predicts human decision time in content moderation
by: Davidson, Thomas
Published: (2025)
by: Davidson, Thomas
Published: (2025)
CAMO: An Agentic Framework for Automated Causal Discovery from Micro Behaviors to Macro Emergence in LLM Agent Simulations
by: Yu, Xiangning, et al.
Published: (2026)
by: Yu, Xiangning, et al.
Published: (2026)
Attributions toward Artificial Agents in a modified Moral Turing Test
by: Aharoni, Eyal, et al.
Published: (2024)
by: Aharoni, Eyal, et al.
Published: (2024)
Causal Abstraction in Model Interpretability: A Compact Survey
by: Zhang, Yihao
Published: (2024)
by: Zhang, Yihao
Published: (2024)
Causal Manifold Fairness: Enforcing Geometric Invariance in Representation Learning
by: Rathore, Vidhi
Published: (2026)
by: Rathore, Vidhi
Published: (2026)
CODE-GEN: A Human-in-the-Loop RAG-Based Agentic AI System for Multiple-Choice Question Generation
by: Duan, Xiaojing, et al.
Published: (2026)
by: Duan, Xiaojing, et al.
Published: (2026)
Invariant Causal Routing for Governing Social Norms in Online Market Economies
by: Yu, Xiangning, et al.
Published: (2026)
by: Yu, Xiangning, et al.
Published: (2026)
Causal Reasoning and Large Language Models: Opening a New Frontier for Causality
by: Kıcıman, Emre, et al.
Published: (2023)
by: Kıcıman, Emre, et al.
Published: (2023)
Teacher agency in the age of generative AI: towards a framework of hybrid intelligence for learning design
by: Frøsig, Thomas B, et al.
Published: (2024)
by: Frøsig, Thomas B, et al.
Published: (2024)
Large Language Models Develop Novel Social Biases Through Adaptive Exploration
by: Wu, Addison J., et al.
Published: (2025)
by: Wu, Addison J., et al.
Published: (2025)
Artificial Intelligence (AI) and the Relationship between Agency, Autonomy, and Moral Patiency
by: Formosa, Paul, et al.
Published: (2025)
by: Formosa, Paul, et al.
Published: (2025)
Similar Items
-
On Probabilistic and Causal Reasoning with Summation Operators
by: Ibeling, Duligur, et al.
Published: (2024) -
How Causal Abstraction Underpins Computational Explanation
by: Geiger, Atticus, et al.
Published: (2025) -
Bucketing the Good Apples: A Method for Diagnosing and Improving Causal Abstraction
by: Puyin, Li, et al.
Published: (2026) -
Belief in the Machine: Investigating Epistemological Blind Spots of Language Models
by: Suzgun, Mirac, et al.
Published: (2024) -
QueerBench: Quantifying Discrimination in Language Models Toward Queer Identities
by: Sosto, Mae, et al.
Published: (2024)