Beyond Model Interpretability: Socio-Structural Explanations in Machine Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Smart, Andrew, Kasirzadeh, Atoosa |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Two Types of AI Existential Risk: Decisive and Accumulative
by: Kasirzadeh, Atoosa
Published: (2024)
by: Kasirzadeh, Atoosa
Published: (2024)
Explanation Hacking: The perils of algorithmic recourse
by: Sullivan, Emily, et al.
Published: (2024)
by: Sullivan, Emily, et al.
Published: (2024)
Position: Beyond Sensitive Attributes, ML Fairness Should Quantify Structural Injustice via Social Determinants
by: Tang, Zeyu, et al.
Published: (2025)
by: Tang, Zeyu, et al.
Published: (2025)
Measurement challenges in AI catastrophic risk governance and safety frameworks
by: Kasirzadeh, Atoosa
Published: (2024)
by: Kasirzadeh, Atoosa
Published: (2024)
SHAPCA: Consistent and Interpretable Explanations for Machine Learning Models on Spectroscopy Data
by: Zhang, Mingxing, et al.
Published: (2026)
by: Zhang, Mingxing, et al.
Published: (2026)
Generative Value Conflicts Reveal LLM Priorities
by: Liu, Andy, et al.
Published: (2025)
by: Liu, Andy, et al.
Published: (2025)
AI Safety for Everyone
by: Gyevnar, Balint, et al.
Published: (2025)
by: Gyevnar, Balint, et al.
Published: (2025)
Characterizing AI Agents for Alignment and Governance
by: Kasirzadeh, Atoosa, et al.
Published: (2025)
by: Kasirzadeh, Atoosa, et al.
Published: (2025)
Bridging the Gap in the Responsible AI Divides
by: Gyevnár, Bálint, et al.
Published: (2026)
by: Gyevnár, Bálint, et al.
Published: (2026)
Beyond Shapley Values: Cooperative Games for the Interpretation of Machine Learning Models
by: Idrissi, Marouane Il, et al.
Published: (2025)
by: Idrissi, Marouane Il, et al.
Published: (2025)
Machine Learning from Explanations
by: Tao, Jiashu, et al.
Published: (2025)
by: Tao, Jiashu, et al.
Published: (2025)
Unified Explanations in Machine Learning Models: A Perturbation Approach
by: Dineen, Jacob, et al.
Published: (2024)
by: Dineen, Jacob, et al.
Published: (2024)
Quantum Spectral Reasoning: A Non-Neural Architecture for Interpretable Machine Learning
by: Kiruluta, Andrew
Published: (2025)
by: Kiruluta, Andrew
Published: (2025)
Interpretable Model-Aware Counterfactual Explanations for Random Forest
by: Harvey, Joshua S., et al.
Published: (2025)
by: Harvey, Joshua S., et al.
Published: (2025)
Linking Model Intervention to Causal Interpretation in Model Explanation
by: Cheng, Debo, et al.
Published: (2024)
by: Cheng, Debo, et al.
Published: (2024)
From Explainability to Interpretability: Interpretable Policies in Reinforcement Learning Via Model Explanation
by: Li, Peilang, et al.
Published: (2025)
by: Li, Peilang, et al.
Published: (2025)
Fast Calibrated Explanations: Efficient and Uncertainty-Aware Explanations for Machine Learning Models
by: Löfström, Tuwe, et al.
Published: (2024)
by: Löfström, Tuwe, et al.
Published: (2024)
Review of Interpretable Machine Learning Models for Disease Prognosis
by: Shen, Jinzhi, et al.
Published: (2024)
by: Shen, Jinzhi, et al.
Published: (2024)
The Effect of Enforcing Fairness on Reshaping Explanations in Machine Learning Models
by: Anderson, Joshua Wolff, et al.
Published: (2025)
by: Anderson, Joshua Wolff, et al.
Published: (2025)
Interpreting Language Reward Models via Contrastive Explanations
by: Jiang, Junqi, et al.
Published: (2024)
by: Jiang, Junqi, et al.
Published: (2024)
Causality-Aware Local Interpretable Model-Agnostic Explanations
by: Cinquini, Martina, et al.
Published: (2022)
by: Cinquini, Martina, et al.
Published: (2022)
Explanation-Guided Adversarial Training for Robust and Interpretable Models
by: Chen, Chao, et al.
Published: (2026)
by: Chen, Chao, et al.
Published: (2026)
Interpreting Graph Inference with Skyline Explanations
by: Qiu, Dazhuo, et al.
Published: (2025)
by: Qiu, Dazhuo, et al.
Published: (2025)
GAMformer: Bridging Tabular Foundation Models and Interpretable Machine Learning
by: Mueller, Andreas, et al.
Published: (2024)
by: Mueller, Andreas, et al.
Published: (2024)
Physics-Inspired Interpretability Of Machine Learning Models
by: Niroomand, Maximilian P, et al.
Published: (2023)
by: Niroomand, Maximilian P, et al.
Published: (2023)
Off-the-Shelf LLMs as Process Scorers: Training-Free Alternative to PRMs for Mathematical Reasoning
by: Chegini, Atoosa, et al.
Published: (2026)
by: Chegini, Atoosa, et al.
Published: (2026)
When Machine Learning Gets Personal: Evaluating Prediction and Explanation
by: Cornelis, Louisa, et al.
Published: (2025)
by: Cornelis, Louisa, et al.
Published: (2025)
A Single Neuron Is Sufficient to Bypass Safety Alignment in Large Language Models
by: Kazemi, Hamid, et al.
Published: (2026)
by: Kazemi, Hamid, et al.
Published: (2026)
PREF-XAI: Preference-Based Personalized Rule Explanations of Black-Box Machine Learning Models
by: Greco, Salvatore, et al.
Published: (2026)
by: Greco, Salvatore, et al.
Published: (2026)
Harnessing Explanations: LLM-to-LM Interpreter for Enhanced Text-Attributed Graph Representation Learning
by: He, Xiaoxin, et al.
Published: (2023)
by: He, Xiaoxin, et al.
Published: (2023)
Beyond Attribution: Unified Concept-Level Explanations
by: Liu, Junhao, et al.
Published: (2024)
by: Liu, Junhao, et al.
Published: (2024)
Policy Trees for Prediction: Interpretable and Adaptive Model Selection for Machine Learning
by: Bertsimas, Dimitris, et al.
Published: (2024)
by: Bertsimas, Dimitris, et al.
Published: (2024)
Interpreting Black-box Machine Learning Models for High Dimensional Datasets
by: Karim, Md. Rezaul, et al.
Published: (2022)
by: Karim, Md. Rezaul, et al.
Published: (2022)
Hyperparameter Optimisation with Practical Interpretability and Explanation Methods in Probabilistic Curriculum Learning
by: Salt, Llewyn, et al.
Published: (2025)
by: Salt, Llewyn, et al.
Published: (2025)
Beyond Euclid: An Illustrated Guide to Modern Machine Learning with Geometric, Topological, and Algebraic Structures
by: Papillon, Mathilde, et al.
Published: (2024)
by: Papillon, Mathilde, et al.
Published: (2024)
Beyond the Black Box: An Interpretable Machine Learning Framework for Predicting Electronic Structure Microdescriptors and Structure-Performance Relationships in Fe-based Catalytic Systems
by: Romiluyi, Oyinkansola
Published: (2026)
by: Romiluyi, Oyinkansola
Published: (2026)
Understanding Disparities in Post Hoc Machine Learning Explanation
by: Mhasawade, Vishwali, et al.
Published: (2024)
by: Mhasawade, Vishwali, et al.
Published: (2024)
Robust Counterfactual Explanations in Machine Learning: A Survey
by: Jiang, Junqi, et al.
Published: (2024)
by: Jiang, Junqi, et al.
Published: (2024)
The Impact of Machine Learning Uncertainty on the Robustness of Counterfactual Explanations
by: Christodoulou, Leonidas, et al.
Published: (2026)
by: Christodoulou, Leonidas, et al.
Published: (2026)
Counterfactual Explanations of Black-box Machine Learning Models using Causal Discovery with Applications to Credit Rating
by: Takahashi, Daisuke, et al.
Published: (2024)
by: Takahashi, Daisuke, et al.
Published: (2024)
Similar Items
-
Two Types of AI Existential Risk: Decisive and Accumulative
by: Kasirzadeh, Atoosa
Published: (2024) -
Explanation Hacking: The perils of algorithmic recourse
by: Sullivan, Emily, et al.
Published: (2024) -
Position: Beyond Sensitive Attributes, ML Fairness Should Quantify Structural Injustice via Social Determinants
by: Tang, Zeyu, et al.
Published: (2025) -
Measurement challenges in AI catastrophic risk governance and safety frameworks
by: Kasirzadeh, Atoosa
Published: (2024) -
SHAPCA: Consistent and Interpretable Explanations for Machine Learning Models on Spectroscopy Data
by: Zhang, Mingxing, et al.
Published: (2026)