Model Science: getting serious about verification, explanation and control of AI systems
Fuente:
arXiv
Saved in:
| Main Authors: | Biecek, Przemyslaw, Samek, Wojciech |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Position: Explain to Question not to Justify
by: Biecek, Przemyslaw, et al.
Published: (2024)
by: Biecek, Przemyslaw, et al.
Published: (2024)
Exploring Local Explanations of Nonlinear Models Using Animated Linear Projections
by: Spyrison, Nicholas, et al.
Published: (2022)
by: Spyrison, Nicholas, et al.
Published: (2022)
The Case for Model Science: Verify, Explore, Steer, Refine
by: Biecek, Przemyslaw, et al.
Published: (2026)
by: Biecek, Przemyslaw, et al.
Published: (2026)
Position: Do Not Explain Vision Models Without Context
by: Tomaszewska, Paulina, et al.
Published: (2024)
by: Tomaszewska, Paulina, et al.
Published: (2024)
Attributions All the Way Down? The Metagame of Interpretability
by: Baniecki, Hubert, et al.
Published: (2026)
by: Baniecki, Hubert, et al.
Published: (2026)
Global Counterfactual Directions
by: Sobieski, Bartlomiej, et al.
Published: (2024)
by: Sobieski, Bartlomiej, et al.
Published: (2024)
Iterative Inference in a Chess-Playing Neural Network
by: Sandmann, Elias, et al.
Published: (2025)
by: Sandmann, Elias, et al.
Published: (2025)
Adversarial attacks and defenses in explainable artificial intelligence: A survey
by: Baniecki, Hubert, et al.
Published: (2023)
by: Baniecki, Hubert, et al.
Published: (2023)
CNN-based explanation ensembling for dataset, representation and explanations evaluation
by: Hryniewska-Guzik, Weronika, et al.
Published: (2024)
by: Hryniewska-Guzik, Weronika, et al.
Published: (2024)
Interpreting CLIP with Hierarchical Sparse Autoencoders
by: Zaigrajew, Vladimir, et al.
Published: (2025)
by: Zaigrajew, Vladimir, et al.
Published: (2025)
Your CLIP has 164 dimensions of noise: Exploring the embeddings covariance eigenspectrum of contrastively pretrained vision-language transformers
by: Grzywaczewski, Jakub, et al.
Published: (2026)
by: Grzywaczewski, Jakub, et al.
Published: (2026)
Ensuring Medical AI Safety: Interpretability-Driven Detection and Mitigation of Spurious Model Behavior and Associated Data
by: Pahde, Frederik, et al.
Published: (2025)
by: Pahde, Frederik, et al.
Published: (2025)
Exploration of the Rashomon Set Assists Trustworthy Explanations for Medical Data
by: Kobylińska, Katarzyna, et al.
Published: (2023)
by: Kobylińska, Katarzyna, et al.
Published: (2023)
Explaining Predictive Uncertainty by Exposing Second-Order Effects
by: Bley, Florian, et al.
Published: (2024)
by: Bley, Florian, et al.
Published: (2024)
Atlas-Alignment: Making Interpretability Transferable Across Language Models
by: Puri, Bruno, et al.
Published: (2025)
by: Puri, Bruno, et al.
Published: (2025)
X-ray transferable polyrepresentation learning
by: Hryniewska-Guzik, Weronika, et al.
Published: (2025)
by: Hryniewska-Guzik, Weronika, et al.
Published: (2025)
NormEnsembleXAI: Unveiling the Strengths and Weaknesses of XAI Ensemble Techniques
by: Hryniewska-Guzik, Weronika, et al.
Published: (2024)
by: Hryniewska-Guzik, Weronika, et al.
Published: (2024)
Mechanistic understanding and validation of large AI models with SemanticLens
by: Dreyer, Maximilian, et al.
Published: (2025)
by: Dreyer, Maximilian, et al.
Published: (2025)
HyConEx: Hypernetwork classifier with counterfactual explanations for tabular data
by: Marszałek, Patryk, et al.
Published: (2025)
by: Marszałek, Patryk, et al.
Published: (2025)
Synthetic Datasets for Machine Learning on Spatio-Temporal Graphs using PDEs
by: Arndt, Jost, et al.
Published: (2025)
by: Arndt, Jost, et al.
Published: (2025)
ECQ$^{\text{x}}$: Explainability-Driven Quantization for Low-Bit and Sparse DNNs
by: Becking, Daniel, et al.
Published: (2021)
by: Becking, Daniel, et al.
Published: (2021)
System-Embedded Diffusion Bridge Models
by: Sobieski, Bartlomiej, et al.
Published: (2025)
by: Sobieski, Bartlomiej, et al.
Published: (2025)
SwordBench: Evaluating Orthogonality of Steering Image Representations
by: Zaigrajew, Vladimir, et al.
Published: (2026)
by: Zaigrajew, Vladimir, et al.
Published: (2026)
LINE: LLM-based Iterative Neuron Explanations for Vision Models
by: Zaigrajew, Vladimir, et al.
Published: (2026)
by: Zaigrajew, Vladimir, et al.
Published: (2026)
Mathematics and Coding are Universal AI Benchmarks
by: Chojecki, Przemyslaw
Published: (2025)
by: Chojecki, Przemyslaw
Published: (2025)
Sparse, Efficient and Explainable Data Attribution with DualXDA
by: Yolcu, Galip Ümit, et al.
Published: (2024)
by: Yolcu, Galip Ümit, et al.
Published: (2024)
Self-Improving AI Agents through Self-Play
by: Chojecki, Przemyslaw
Published: (2025)
by: Chojecki, Przemyslaw
Published: (2025)
survex: an R package for explaining machine learning survival models
by: Spytek, Mikołaj, et al.
Published: (2023)
by: Spytek, Mikołaj, et al.
Published: (2023)
$α$-TCAV: A Unified Framework for Testing with Concept Activation Vectors
by: Schnoor, Ekkehard, et al.
Published: (2026)
by: Schnoor, Ekkehard, et al.
Published: (2026)
From What to How: Attributing CLIP's Latent Components Reveals Unexpected Semantic Reliance
by: Dreyer, Maximilian, et al.
Published: (2025)
by: Dreyer, Maximilian, et al.
Published: (2025)
The Clever Hans Effect in Unsupervised Learning
by: Kauffmann, Jacob, et al.
Published: (2024)
by: Kauffmann, Jacob, et al.
Published: (2024)
Post-Hoc Concept Disentanglement: From Correlated to Isolated Concept Representations
by: Erogullari, Eren, et al.
Published: (2025)
by: Erogullari, Eren, et al.
Published: (2025)
Optimizing Federated Learning by Entropy-Based Client Selection
by: Lutz, Andreas, et al.
Published: (2024)
by: Lutz, Andreas, et al.
Published: (2024)
Reactive Model Correction: Mitigating Harm to Task-Relevant Features via Conditional Bias Suppression
by: Bareeva, Dilyara, et al.
Published: (2024)
by: Bareeva, Dilyara, et al.
Published: (2024)
From Attribution Maps to Human-Understandable Explanations through Concept Relevance Propagation
by: Achtibat, Reduan, et al.
Published: (2022)
by: Achtibat, Reduan, et al.
Published: (2022)
Towards a perturbation-based explanation for medical AI as differentiable programs
by: Abe, Takeshi, et al.
Published: (2025)
by: Abe, Takeshi, et al.
Published: (2025)
Psychometric Tests for AI Agents and Their Moduli Space
by: Chojecki, Przemyslaw
Published: (2025)
by: Chojecki, Przemyslaw
Published: (2025)
Red-Teaming Segment Anything Model
by: Jankowski, Krzysztof, et al.
Published: (2024)
by: Jankowski, Krzysztof, et al.
Published: (2024)
Relevance-driven Input Dropout: an Explanation-guided Regularization Technique
by: Gururaj, Shreyas, et al.
Published: (2025)
by: Gururaj, Shreyas, et al.
Published: (2025)
Quanda: An Interpretability Toolkit for Training Data Attribution Evaluation and Beyond
by: Bareeva, Dilyara, et al.
Published: (2024)
by: Bareeva, Dilyara, et al.
Published: (2024)
Similar Items
-
Position: Explain to Question not to Justify
by: Biecek, Przemyslaw, et al.
Published: (2024) -
Exploring Local Explanations of Nonlinear Models Using Animated Linear Projections
by: Spyrison, Nicholas, et al.
Published: (2022) -
The Case for Model Science: Verify, Explore, Steer, Refine
by: Biecek, Przemyslaw, et al.
Published: (2026) -
Position: Do Not Explain Vision Models Without Context
by: Tomaszewska, Paulina, et al.
Published: (2024) -
Attributions All the Way Down? The Metagame of Interpretability
by: Baniecki, Hubert, et al.
Published: (2026)