Black Box Model Explanations and the Human Interpretability Expectations -- An Analysis in the Context of Homicide Prediction
Fuente:
arXiv
Saved in:
| Main Authors: | Ribeiro, José, Carneiro, Níkolas, Alves, Ronnie |
|---|---|
| Format: | Preprint |
| Published: |
2022
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
How Reliable and Stable are Explanations of XAI Methods?
by: Ribeiro, José, et al.
Published: (2024)
by: Ribeiro, José, et al.
Published: (2024)
Explanations Based on Item Response Theory (eXirt): A Model-Specific Method to Explain Tree-Ensemble Model in Trust Perspective
by: Ribeiro, José, et al.
Published: (2022)
by: Ribeiro, José, et al.
Published: (2022)
Evaluating Model Explanations without Ground Truth
by: Rawal, Kaivalya, et al.
Published: (2025)
by: Rawal, Kaivalya, et al.
Published: (2025)
Neural Reasoning Networks: Efficient Interpretable Neural Networks With Automatic Textual Explanations
by: Carrow, Stephen, et al.
Published: (2024)
by: Carrow, Stephen, et al.
Published: (2024)
Deep Memory Search: A Metaheuristic Approach for Optimizing Heuristic Search
by: Hedar, Abdel-Rahman, et al.
Published: (2024)
by: Hedar, Abdel-Rahman, et al.
Published: (2024)
Superposition Is Not Necessary: A Mechanistic Interpretability Analysis of Transformer Representations for Time Series Forecasting
by: Yıldırım, Alper
Published: (2026)
by: Yıldırım, Alper
Published: (2026)
Relevance-driven Input Dropout: an Explanation-guided Regularization Technique
by: Gururaj, Shreyas, et al.
Published: (2025)
by: Gururaj, Shreyas, et al.
Published: (2025)
Logic-based Explanations for Linear Support Vector Classifiers with Reject Option
by: Filho, Francisco Mateus Rocha, et al.
Published: (2024)
by: Filho, Francisco Mateus Rocha, et al.
Published: (2024)
FORCE: Feature-Oriented Representation with Clustering and Explanation
by: Mukherjee, Rishav, et al.
Published: (2025)
by: Mukherjee, Rishav, et al.
Published: (2025)
A Simple Generalisation of the Implicit Dynamics of In-Context Learning
by: Innocenti, Francesco, et al.
Published: (2025)
by: Innocenti, Francesco, et al.
Published: (2025)
Interpretability-Guided Bi-objective Optimization: Aligning Accuracy and Explainability
by: Fouladi, Kasra, et al.
Published: (2026)
by: Fouladi, Kasra, et al.
Published: (2026)
An Incremental MaxSAT-based Model to Learn Interpretable and Balanced Classification Rules
by: Júnior, Antônio Carlos Souza Ferreira, et al.
Published: (2024)
by: Júnior, Antônio Carlos Souza Ferreira, et al.
Published: (2024)
Enhancing Classifier Evaluation: A Fairer Benchmarking Strategy Based on Ability and Robustness
by: Cardoso, Lucas, et al.
Published: (2025)
by: Cardoso, Lucas, et al.
Published: (2025)
Fusion-Based Neural Generalization for Predicting Temperature Fields in Industrial PET Preform Heating
by: Alsheikh, Ahmad, et al.
Published: (2025)
by: Alsheikh, Ahmad, et al.
Published: (2025)
Beyond Random Sampling: Instance Quality-Based Data Partitioning via Item Response Theory
by: Cardoso, Lucas, et al.
Published: (2025)
by: Cardoso, Lucas, et al.
Published: (2025)
Manipulating Predictions over Discrete Inputs in Machine Teaching
by: Wu, Xiaodong, et al.
Published: (2024)
by: Wu, Xiaodong, et al.
Published: (2024)
LLM-Driven Intrinsic Motivation for Sparse Reward Reinforcement Learning
by: Quadros, André, et al.
Published: (2025)
by: Quadros, André, et al.
Published: (2025)
STACHE: Local Black-Box Explanations for Reinforcement Learning Policies
by: Elashkin, Andrew, et al.
Published: (2025)
by: Elashkin, Andrew, et al.
Published: (2025)
On the Role of Pre-trained Embeddings in Binary Code Analysis
by: Maier, Alwin, et al.
Published: (2025)
by: Maier, Alwin, et al.
Published: (2025)
Prototype Transformer: Towards Language Model Architectures Interpretable by Design
by: Yordanov, Yordan, et al.
Published: (2026)
by: Yordanov, Yordan, et al.
Published: (2026)
Simulation-Driven Railway Delay Prediction: An Imitation Learning Approach
by: Elliker, Clément, et al.
Published: (2025)
by: Elliker, Clément, et al.
Published: (2025)
Entropy-Based Measurement of Value Drift and Alignment Work in Large Language Models
by: Fadli, Samih
Published: (2025)
by: Fadli, Samih
Published: (2025)
CircuitProbe: Predicting Reasoning Circuits in Transformers via Stability Zone Detection
by: Panuganti, Rajkiran
Published: (2026)
by: Panuganti, Rajkiran
Published: (2026)
First-Mover Bias in Gradient Boosting Explanations: Mechanism, Detection, and Resolution
by: Caraker, Drake, et al.
Published: (2026)
by: Caraker, Drake, et al.
Published: (2026)
Large Language Models as Attribution Regularizers for Efficient Model Training
by: Vukadin, Davor, et al.
Published: (2025)
by: Vukadin, Davor, et al.
Published: (2025)
Compositional Concept-Based Neuron-Level Interpretability for Deep Reinforcement Learning
by: Jiang, Zeyu, et al.
Published: (2025)
by: Jiang, Zeyu, et al.
Published: (2025)
SMOSE: Sparse Mixture of Shallow Experts for Interpretable Reinforcement Learning in Continuous Control Tasks
by: Vincze, Mátyás, et al.
Published: (2024)
by: Vincze, Mátyás, et al.
Published: (2024)
BOND: License to Train with Black-Box Functions
by: Clark, Andrew, et al.
Published: (2025)
by: Clark, Andrew, et al.
Published: (2025)
Extreme AutoML: Analysis of Classification, Regression, and NLP Performance
by: Ratner, Edward, et al.
Published: (2024)
by: Ratner, Edward, et al.
Published: (2024)
Social Interpretable Reinforcement Learning
by: Custode, Leonardo Lucio, et al.
Published: (2024)
by: Custode, Leonardo Lucio, et al.
Published: (2024)
Approximate Domain Unlearning for Vision-Language Models
by: Kawamura, Kodai, et al.
Published: (2025)
by: Kawamura, Kodai, et al.
Published: (2025)
Grokking Beyond the Euclidean Norm of Model Parameters
by: Notsawo, Pascal Jr Tikeng, et al.
Published: (2025)
by: Notsawo, Pascal Jr Tikeng, et al.
Published: (2025)
An Interpretable Rule Creation Method for Black-Box Models based on Surrogate Trees -- SRules
by: Verdasco, Mario Parrón, et al.
Published: (2024)
by: Verdasco, Mario Parrón, et al.
Published: (2024)
Scaling Offline RL via Efficient and Expressive Shortcut Models
by: Espinosa-Dice, Nicolas, et al.
Published: (2025)
by: Espinosa-Dice, Nicolas, et al.
Published: (2025)
DrugReasoner: Interpretable Drug Approval Prediction with a Reasoning-augmented Language Model
by: Ghaffarzadeh-Esfahani, Mohammadreza, et al.
Published: (2025)
by: Ghaffarzadeh-Esfahani, Mohammadreza, et al.
Published: (2025)
RobustBlack: Challenging Black-Box Adversarial Attacks on State-of-the-Art Defenses
by: Djilani, Mohamed, et al.
Published: (2024)
by: Djilani, Mohamed, et al.
Published: (2024)
Architectural Proprioception in State Space Models: Thermodynamic Training Induces Anticipatory Halt Detection
by: Noon, Jay
Published: (2026)
by: Noon, Jay
Published: (2026)
A Practical Approach to using Supervised Machine Learning Models to Classify Aviation Safety Occurrences
by: Siow, Bryan Y.
Published: (2025)
by: Siow, Bryan Y.
Published: (2025)
Model Fusion via Retrofitting
by: Luenam, Phoomraphee, et al.
Published: (2025)
by: Luenam, Phoomraphee, et al.
Published: (2025)
Causally Grounded Mechanistic Interpretability for LLMs with Faithful Natural-Language Explanations
by: Mahale, Ajay Pravin
Published: (2026)
by: Mahale, Ajay Pravin
Published: (2026)
Similar Items
-
How Reliable and Stable are Explanations of XAI Methods?
by: Ribeiro, José, et al.
Published: (2024) -
Explanations Based on Item Response Theory (eXirt): A Model-Specific Method to Explain Tree-Ensemble Model in Trust Perspective
by: Ribeiro, José, et al.
Published: (2022) -
Evaluating Model Explanations without Ground Truth
by: Rawal, Kaivalya, et al.
Published: (2025) -
Neural Reasoning Networks: Efficient Interpretable Neural Networks With Automatic Textual Explanations
by: Carrow, Stephen, et al.
Published: (2024) -
Deep Memory Search: A Metaheuristic Approach for Optimizing Heuristic Search
by: Hedar, Abdel-Rahman, et al.
Published: (2024)