PURE: Turning Polysemantic Neurons Into Pure Features by Identifying Relevant Circuits
Fuente:
arXiv
Guardado en:
| Autores principales: | Dreyer, Maximilian, Purelku, Erblina, Vielhaben, Johanna, Samek, Wojciech, Lapuschkin, Sebastian |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Dyslexify: A Mechanistic Defense Against Typographic Attacks in CLIP
por: Hufe, Lorenz, et al.
Publicado: (2025)
por: Hufe, Lorenz, et al.
Publicado: (2025)
Reactive Model Correction: Mitigating Harm to Task-Relevant Features via Conditional Bias Suppression
por: Bareeva, Dilyara, et al.
Publicado: (2024)
por: Bareeva, Dilyara, et al.
Publicado: (2024)
Understanding the (Extra-)Ordinary: Validating Deep Model Decisions with Prototypical Concept-based Explanations
por: Dreyer, Maximilian, et al.
Publicado: (2023)
por: Dreyer, Maximilian, et al.
Publicado: (2023)
Contrastive Semantic Projection: Faithful Neuron Labeling with Contrastive Examples
por: Bouanani, Oussama, et al.
Publicado: (2026)
por: Bouanani, Oussama, et al.
Publicado: (2026)
AttnLRP: Attention-Aware Layer-Wise Relevance Propagation for Transformers
por: Achtibat, Reduan, et al.
Publicado: (2024)
por: Achtibat, Reduan, et al.
Publicado: (2024)
Beyond Scalars: Concept-Based Alignment Analysis in Vision Transformers
por: Vielhaben, Johanna, et al.
Publicado: (2024)
por: Vielhaben, Johanna, et al.
Publicado: (2024)
Pruning By Explaining Revisited: Optimizing Attribution Methods to Prune CNNs and Transformers
por: Hatefi, Sayed Mohammad Vakilzadeh, et al.
Publicado: (2024)
por: Hatefi, Sayed Mohammad Vakilzadeh, et al.
Publicado: (2024)
Post-Hoc Concept Disentanglement: From Correlated to Isolated Concept Representations
por: Erogullari, Eren, et al.
Publicado: (2025)
por: Erogullari, Eren, et al.
Publicado: (2025)
Ensuring Medical AI Safety: Interpretability-Driven Detection and Mitigation of Spurious Model Behavior and Associated Data
por: Pahde, Frederik, et al.
Publicado: (2025)
por: Pahde, Frederik, et al.
Publicado: (2025)
Navigating Neural Space: Revisiting Concept Activation Vectors to Overcome Directional Divergence
por: Pahde, Frederik, et al.
Publicado: (2022)
por: Pahde, Frederik, et al.
Publicado: (2022)
Human-Centered Evaluation of XAI Methods
por: Dawoud, Karam, et al.
Publicado: (2023)
por: Dawoud, Karam, et al.
Publicado: (2023)
Mechanistic understanding and validation of large AI models with SemanticLens
por: Dreyer, Maximilian, et al.
Publicado: (2025)
por: Dreyer, Maximilian, et al.
Publicado: (2025)
Explainable concept mappings of MRI: Revealing the mechanisms underlying deep learning-based brain disease classification
por: Tinauer, Christian, et al.
Publicado: (2024)
por: Tinauer, Christian, et al.
Publicado: (2024)
Decoupling Pixel Flipping and Occlusion Strategy for Consistent XAI Benchmarks
por: Blücher, Stefan, et al.
Publicado: (2024)
por: Blücher, Stefan, et al.
Publicado: (2024)
FeatInv: Spatially resolved mapping from feature space to input space using conditional diffusion models
por: Neukirch, Nils, et al.
Publicado: (2025)
por: Neukirch, Nils, et al.
Publicado: (2025)
SCAM: A Real-World Typographic Robustness Evaluation for Multimodal Foundation Models
por: Westerhoff, Justus, et al.
Publicado: (2025)
por: Westerhoff, Justus, et al.
Publicado: (2025)
From Attribution Maps to Human-Understandable Explanations through Concept Relevance Propagation
por: Achtibat, Reduan, et al.
Publicado: (2022)
por: Achtibat, Reduan, et al.
Publicado: (2022)
ECQ$^{\text{x}}$: Explainability-Driven Quantization for Low-Bit and Sparse DNNs
por: Becking, Daniel, et al.
Publicado: (2021)
por: Becking, Daniel, et al.
Publicado: (2021)
XAI-guided Insulator Anomaly Detection for Imbalanced Datasets
por: Hoefler, Maximilian Andreas, et al.
Publicado: (2024)
por: Hoefler, Maximilian Andreas, et al.
Publicado: (2024)
From What to How: Attributing CLIP's Latent Components Reveals Unexpected Semantic Reliance
por: Dreyer, Maximilian, et al.
Publicado: (2025)
por: Dreyer, Maximilian, et al.
Publicado: (2025)
Playing the network backward: A Game Theoretic Attribution Framework
por: Zimmermann, Jakob Paul, et al.
Publicado: (2026)
por: Zimmermann, Jakob Paul, et al.
Publicado: (2026)
Iterative Inference in a Chess-Playing Neural Network
por: Sandmann, Elias, et al.
Publicado: (2025)
por: Sandmann, Elias, et al.
Publicado: (2025)
CoE: Chain-of-Explanation via Automatic Visual Concept Circuit Description and Polysemanticity Quantification
por: Yu, Wenlong, et al.
Publicado: (2025)
por: Yu, Wenlong, et al.
Publicado: (2025)
A Fresh Look at Sanity Checks for Saliency Maps
por: Hedström, Anna, et al.
Publicado: (2024)
por: Hedström, Anna, et al.
Publicado: (2024)
Deep Learning-based Multi Project InP Wafer Simulation for Unsupervised Surface Defect Detection
por: Cantú, Emílio Dolgener, et al.
Publicado: (2025)
por: Cantú, Emílio Dolgener, et al.
Publicado: (2025)
Disentangling Polysemantic Neurons with a Null-Calibrated Polysemanticity Index and Causal Patch Interventions
por: Gupta, Manan, et al.
Publicado: (2025)
por: Gupta, Manan, et al.
Publicado: (2025)
Manipulating Feature Visualizations with Gradient Slingshots
por: Bareeva, Dilyara, et al.
Publicado: (2024)
por: Bareeva, Dilyara, et al.
Publicado: (2024)
Fractional Diffusion Bridge Models
por: Nobis, Gabriel, et al.
Publicado: (2025)
por: Nobis, Gabriel, et al.
Publicado: (2025)
Relevance-driven Input Dropout: an Explanation-guided Regularization Technique
por: Gururaj, Shreyas, et al.
Publicado: (2025)
por: Gururaj, Shreyas, et al.
Publicado: (2025)
Attribution-Guided Pruning for Insight and Control: Circuit Discovery and Targeted Correction in Small-scale LLMs
por: Hatefi, Sayed Mohammad Vakilzadeh, et al.
Publicado: (2025)
por: Hatefi, Sayed Mohammad Vakilzadeh, et al.
Publicado: (2025)
Explaining Predictive Uncertainty by Exposing Second-Order Effects
por: Bley, Florian, et al.
Publicado: (2024)
por: Bley, Florian, et al.
Publicado: (2024)
From Attribution to Action: A Human-Centered Application of Activation Steering
por: Labarta, Tobias, et al.
Publicado: (2026)
por: Labarta, Tobias, et al.
Publicado: (2026)
Circuit Insights: Towards Interpretability Beyond Activations
por: Golimblevskaia, Elena, et al.
Publicado: (2025)
por: Golimblevskaia, Elena, et al.
Publicado: (2025)
Steering CLIP's vision transformer with sparse autoencoders
por: Joseph, Sonia, et al.
Publicado: (2025)
por: Joseph, Sonia, et al.
Publicado: (2025)
Learning Robust Convolutional Neural Networks with Relevant Feature Focusing via Explanations
por: Adachi, Kazuki, et al.
Publicado: (2022)
por: Adachi, Kazuki, et al.
Publicado: (2022)
DiG-IN: Diffusion Guidance for Investigating Networks -- Uncovering Classifier Differences Neuron Visualisations and Visual Counterfactual Explanations
por: Augustin, Maximilian, et al.
Publicado: (2023)
por: Augustin, Maximilian, et al.
Publicado: (2023)
Synthetic Generation of Dermatoscopic Images with GAN and Closed-Form Factorization
por: Mekala, Rohan Reddy, et al.
Publicado: (2024)
por: Mekala, Rohan Reddy, et al.
Publicado: (2024)
Atlas-Alignment: Making Interpretability Transferable Across Language Models
por: Puri, Bruno, et al.
Publicado: (2025)
por: Puri, Bruno, et al.
Publicado: (2025)
Prisma: An Open Source Toolkit for Mechanistic Interpretability in Vision and Video
por: Joseph, Sonia, et al.
Publicado: (2025)
por: Joseph, Sonia, et al.
Publicado: (2025)
CLIP-PAE: Projection-Augmentation Embedding to Extract Relevant Features for a Disentangled, Interpretable, and Controllable Text-Guided Face Manipulation
por: Zhou, Chenliang, et al.
Publicado: (2022)
por: Zhou, Chenliang, et al.
Publicado: (2022)
Ejemplares similares
-
Dyslexify: A Mechanistic Defense Against Typographic Attacks in CLIP
por: Hufe, Lorenz, et al.
Publicado: (2025) -
Reactive Model Correction: Mitigating Harm to Task-Relevant Features via Conditional Bias Suppression
por: Bareeva, Dilyara, et al.
Publicado: (2024) -
Understanding the (Extra-)Ordinary: Validating Deep Model Decisions with Prototypical Concept-based Explanations
por: Dreyer, Maximilian, et al.
Publicado: (2023) -
Contrastive Semantic Projection: Faithful Neuron Labeling with Contrastive Examples
por: Bouanani, Oussama, et al.
Publicado: (2026) -
AttnLRP: Attention-Aware Layer-Wise Relevance Propagation for Transformers
por: Achtibat, Reduan, et al.
Publicado: (2024)