The Case for Model Science: Verify, Explore, Steer, Refine
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Biecek, Przemyslaw, Longo, Luca, Zhou, Jianlong, Fel, Thomas, Holzinger, Andreas, Samek, Wojciech |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Model Science: getting serious about verification, explanation and control of AI systems
von: Biecek, Przemyslaw, et al.
Veröffentlicht: (2025)
von: Biecek, Przemyslaw, et al.
Veröffentlicht: (2025)
Position: Explain to Question not to Justify
von: Biecek, Przemyslaw, et al.
Veröffentlicht: (2024)
von: Biecek, Przemyslaw, et al.
Veröffentlicht: (2024)
CNN-based explanation ensembling for dataset, representation and explanations evaluation
von: Hryniewska-Guzik, Weronika, et al.
Veröffentlicht: (2024)
von: Hryniewska-Guzik, Weronika, et al.
Veröffentlicht: (2024)
Exploring Local Explanations of Nonlinear Models Using Animated Linear Projections
von: Spyrison, Nicholas, et al.
Veröffentlicht: (2022)
von: Spyrison, Nicholas, et al.
Veröffentlicht: (2022)
SwordBench: Evaluating Orthogonality of Steering Image Representations
von: Zaigrajew, Vladimir, et al.
Veröffentlicht: (2026)
von: Zaigrajew, Vladimir, et al.
Veröffentlicht: (2026)
Ethical ChatGPT: Concerns, Challenges, and Commandments
von: Zhou, Jianlong, et al.
Veröffentlicht: (2023)
von: Zhou, Jianlong, et al.
Veröffentlicht: (2023)
Position: Do Not Explain Vision Models Without Context
von: Tomaszewska, Paulina, et al.
Veröffentlicht: (2024)
von: Tomaszewska, Paulina, et al.
Veröffentlicht: (2024)
Your CLIP has 164 dimensions of noise: Exploring the embeddings covariance eigenspectrum of contrastively pretrained vision-language transformers
von: Grzywaczewski, Jakub, et al.
Veröffentlicht: (2026)
von: Grzywaczewski, Jakub, et al.
Veröffentlicht: (2026)
Global Counterfactual Directions
von: Sobieski, Bartlomiej, et al.
Veröffentlicht: (2024)
von: Sobieski, Bartlomiej, et al.
Veröffentlicht: (2024)
Adversarial attacks and defenses in explainable artificial intelligence: A survey
von: Baniecki, Hubert, et al.
Veröffentlicht: (2023)
von: Baniecki, Hubert, et al.
Veröffentlicht: (2023)
Attributions All the Way Down? The Metagame of Interpretability
von: Baniecki, Hubert, et al.
Veröffentlicht: (2026)
von: Baniecki, Hubert, et al.
Veröffentlicht: (2026)
Sparks of Explainability: Recent Advancements in Explaining Large Vision Models
von: Fel, Thomas
Veröffentlicht: (2025)
von: Fel, Thomas
Veröffentlicht: (2025)
XAI-guided Insulator Anomaly Detection for Imbalanced Datasets
von: Hoefler, Maximilian Andreas, et al.
Veröffentlicht: (2024)
von: Hoefler, Maximilian Andreas, et al.
Veröffentlicht: (2024)
X-ray transferable polyrepresentation learning
von: Hryniewska-Guzik, Weronika, et al.
Veröffentlicht: (2025)
von: Hryniewska-Guzik, Weronika, et al.
Veröffentlicht: (2025)
Interpreting CLIP with Hierarchical Sparse Autoencoders
von: Zaigrajew, Vladimir, et al.
Veröffentlicht: (2025)
von: Zaigrajew, Vladimir, et al.
Veröffentlicht: (2025)
Iterative Inference in a Chess-Playing Neural Network
von: Sandmann, Elias, et al.
Veröffentlicht: (2025)
von: Sandmann, Elias, et al.
Veröffentlicht: (2025)
Ensuring Medical AI Safety: Interpretability-Driven Detection and Mitigation of Spurious Model Behavior and Associated Data
von: Pahde, Frederik, et al.
Veröffentlicht: (2025)
von: Pahde, Frederik, et al.
Veröffentlicht: (2025)
NormEnsembleXAI: Unveiling the Strengths and Weaknesses of XAI Ensemble Techniques
von: Hryniewska-Guzik, Weronika, et al.
Veröffentlicht: (2024)
von: Hryniewska-Guzik, Weronika, et al.
Veröffentlicht: (2024)
Optimizing Federated Learning by Entropy-Based Client Selection
von: Lutz, Andreas, et al.
Veröffentlicht: (2024)
von: Lutz, Andreas, et al.
Veröffentlicht: (2024)
Atlas-Alignment: Making Interpretability Transferable Across Language Models
von: Puri, Bruno, et al.
Veröffentlicht: (2025)
von: Puri, Bruno, et al.
Veröffentlicht: (2025)
Understanding the (Extra-)Ordinary: Validating Deep Model Decisions with Prototypical Concept-based Explanations
von: Dreyer, Maximilian, et al.
Veröffentlicht: (2023)
von: Dreyer, Maximilian, et al.
Veröffentlicht: (2023)
Exploration of the Rashomon Set Assists Trustworthy Explanations for Medical Data
von: Kobylińska, Katarzyna, et al.
Veröffentlicht: (2023)
von: Kobylińska, Katarzyna, et al.
Veröffentlicht: (2023)
Explaining Predictive Uncertainty by Exposing Second-Order Effects
von: Bley, Florian, et al.
Veröffentlicht: (2024)
von: Bley, Florian, et al.
Veröffentlicht: (2024)
LINE: LLM-based Iterative Neuron Explanations for Vision Models
von: Zaigrajew, Vladimir, et al.
Veröffentlicht: (2026)
von: Zaigrajew, Vladimir, et al.
Veröffentlicht: (2026)
Steering CLIP's vision transformer with sparse autoencoders
von: Joseph, Sonia, et al.
Veröffentlicht: (2025)
von: Joseph, Sonia, et al.
Veröffentlicht: (2025)
Sparse, Efficient and Explainable Data Attribution with DualXDA
von: Yolcu, Galip Ümit, et al.
Veröffentlicht: (2024)
von: Yolcu, Galip Ümit, et al.
Veröffentlicht: (2024)
From Attribution to Action: A Human-Centered Application of Activation Steering
von: Labarta, Tobias, et al.
Veröffentlicht: (2026)
von: Labarta, Tobias, et al.
Veröffentlicht: (2026)
The Dark Patterns of Personalized Persuasion in Large Language Models: Exposing Persuasive Linguistic Features for Big Five Personality Traits in LLMs Responses
von: Mieleszczenko-Kowszewicz, Wiktoria, et al.
Veröffentlicht: (2024)
von: Mieleszczenko-Kowszewicz, Wiktoria, et al.
Veröffentlicht: (2024)
Rethinking Visual Counterfactual Explanations Through Region Constraint
von: Sobieski, Bartlomiej, et al.
Veröffentlicht: (2024)
von: Sobieski, Bartlomiej, et al.
Veröffentlicht: (2024)
Structural Compactness as a Complementary Criterion for Explanation Quality
von: Mesgari, Mohammad Mahdi, et al.
Veröffentlicht: (2026)
von: Mesgari, Mohammad Mahdi, et al.
Veröffentlicht: (2026)
From What to How: Attributing CLIP's Latent Components Reveals Unexpected Semantic Reliance
von: Dreyer, Maximilian, et al.
Veröffentlicht: (2025)
von: Dreyer, Maximilian, et al.
Veröffentlicht: (2025)
The System Hallucination Scale (SHS): A Minimal yet Effective Human-Centered Instrument for Evaluating Hallucination-Related Behavior in Large Language Models
von: Müller, Heimo, et al.
Veröffentlicht: (2026)
von: Müller, Heimo, et al.
Veröffentlicht: (2026)
Synthetic Datasets for Machine Learning on Spatio-Temporal Graphs using PDEs
von: Arndt, Jost, et al.
Veröffentlicht: (2025)
von: Arndt, Jost, et al.
Veröffentlicht: (2025)
ECQ$^{\text{x}}$: Explainability-Driven Quantization for Low-Bit and Sparse DNNs
von: Becking, Daniel, et al.
Veröffentlicht: (2021)
von: Becking, Daniel, et al.
Veröffentlicht: (2021)
System-Embedded Diffusion Bridge Models
von: Sobieski, Bartlomiej, et al.
Veröffentlicht: (2025)
von: Sobieski, Bartlomiej, et al.
Veröffentlicht: (2025)
Post-Hoc Concept Disentanglement: From Correlated to Isolated Concept Representations
von: Erogullari, Eren, et al.
Veröffentlicht: (2025)
von: Erogullari, Eren, et al.
Veröffentlicht: (2025)
Mind What You Ask For: Emotional and Rational Faces of Persuasion by Large Language Models
von: Mieleszczenko-Kowszewicz, Wiktoria, et al.
Veröffentlicht: (2025)
von: Mieleszczenko-Kowszewicz, Wiktoria, et al.
Veröffentlicht: (2025)
Reactive Model Correction: Mitigating Harm to Task-Relevant Features via Conditional Bias Suppression
von: Bareeva, Dilyara, et al.
Veröffentlicht: (2024)
von: Bareeva, Dilyara, et al.
Veröffentlicht: (2024)
$α$-TCAV: A Unified Framework for Testing with Concept Activation Vectors
von: Schnoor, Ekkehard, et al.
Veröffentlicht: (2026)
von: Schnoor, Ekkehard, et al.
Veröffentlicht: (2026)
Local Intrinsic Dimension Unveils Hallucinations in Diffusion Models
von: Sobieski, Bartlomiej, et al.
Veröffentlicht: (2026)
von: Sobieski, Bartlomiej, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Model Science: getting serious about verification, explanation and control of AI systems
von: Biecek, Przemyslaw, et al.
Veröffentlicht: (2025) -
Position: Explain to Question not to Justify
von: Biecek, Przemyslaw, et al.
Veröffentlicht: (2024) -
CNN-based explanation ensembling for dataset, representation and explanations evaluation
von: Hryniewska-Guzik, Weronika, et al.
Veröffentlicht: (2024) -
Exploring Local Explanations of Nonlinear Models Using Animated Linear Projections
von: Spyrison, Nicholas, et al.
Veröffentlicht: (2022) -
SwordBench: Evaluating Orthogonality of Steering Image Representations
von: Zaigrajew, Vladimir, et al.
Veröffentlicht: (2026)