Mechanistic Interpretability for AI Safety -- A Review
Fuente:
arXiv
Saved in:
| Main Authors: | Bereska, Leonard, Gavves, Efstratios |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Superposition as Lossy Compression: Measure with Sparse Autoencoders and Connect to Adversarial Vulnerability
by: Bereska, Leonard, et al.
Published: (2025)
by: Bereska, Leonard, et al.
Published: (2025)
Mechanistic Neural Networks for Scientific Machine Learning
by: Pervez, Adeel, et al.
Published: (2024)
by: Pervez, Adeel, et al.
Published: (2024)
From MLP to NeoMLP: Leveraging Self-Attention for Neural Fields
by: Kofinas, Miltiadis, et al.
Published: (2024)
by: Kofinas, Miltiadis, et al.
Published: (2024)
Physics-Guided Radiotherapy Treatment Planning with Deep Learning
by: Achlatis, Stefanos, et al.
Published: (2025)
by: Achlatis, Stefanos, et al.
Published: (2025)
Unleashing Uncertainty: Efficient Machine Unlearning for Generative AI
by: Spartalis, Christoforos N., et al.
Published: (2025)
by: Spartalis, Christoforos N., et al.
Published: (2025)
From Explainable to Explained AI: Ideas for Falsifying and Quantifying Explanations
by: Schirris, Yoni, et al.
Published: (2025)
by: Schirris, Yoni, et al.
Published: (2025)
Language Agents Meet Causality -- Bridging LLMs and Causal World Models
by: Gkountouras, John, et al.
Published: (2024)
by: Gkountouras, John, et al.
Published: (2024)
LoTUS: Large-Scale Machine Unlearning with a Taste of Uncertainty
by: Spartalis, Christoforos N., et al.
Published: (2025)
by: Spartalis, Christoforos N., et al.
Published: (2025)
Mechanistic PDE Networks for Discovery of Governing Equations
by: Pervez, Adeel, et al.
Published: (2025)
by: Pervez, Adeel, et al.
Published: (2025)
CTRL-O: Language-Controllable Object-Centric Visual Representation Learning
by: Didolkar, Aniket, et al.
Published: (2025)
by: Didolkar, Aniket, et al.
Published: (2025)
Distributional Vision-Language Alignment by Cauchy-Schwarz Divergence
by: Yin, Wenzhe, et al.
Published: (2025)
by: Yin, Wenzhe, et al.
Published: (2025)
Towards Unified Attribution in Explainable AI, Data-Centric AI, and Mechanistic Interpretability
by: Zhang, Shichang, et al.
Published: (2025)
by: Zhang, Shichang, et al.
Published: (2025)
AI Alignment Strategies from a Risk Perspective: Independent Safety Mechanisms or Shared Failures?
by: Dung, Leonard, et al.
Published: (2025)
by: Dung, Leonard, et al.
Published: (2025)
Mechanistic Interpretability of LoRA-Adapted Language Models for Nuclear Reactor Safety Applications
by: Lee, Yoon Pyo
Published: (2025)
by: Lee, Yoon Pyo
Published: (2025)
Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability
by: Geiger, Atticus, et al.
Published: (2023)
by: Geiger, Atticus, et al.
Published: (2023)
Mechanistic Interpretability Needs Philosophy
by: Williams, Iwan, et al.
Published: (2025)
by: Williams, Iwan, et al.
Published: (2025)
Mechanistically Interpreting Compression in Vision-Language Models
by: Elluru, Veeraraju, et al.
Published: (2026)
by: Elluru, Veeraraju, et al.
Published: (2026)
Safety by Measurement: A Systematic Literature Review of AI Safety Evaluation Methods
by: Grey, Markov, et al.
Published: (2025)
by: Grey, Markov, et al.
Published: (2025)
A Practical Review of Mechanistic Interpretability for Transformer-Based Language Models
by: Rai, Daking, et al.
Published: (2024)
by: Rai, Daking, et al.
Published: (2024)
Challenges in Mechanistically Interpreting Model Representations
by: Golechha, Satvik, et al.
Published: (2024)
by: Golechha, Satvik, et al.
Published: (2024)
Mechanistic Interpretability in the Presence of Architectural Obfuscation
by: Florencio, Marcos, et al.
Published: (2025)
by: Florencio, Marcos, et al.
Published: (2025)
Neural Transparency: Mechanistic Interpretability Interfaces for Anticipating Model Behaviors for Personalized AI
by: Karny, Sheer, et al.
Published: (2025)
by: Karny, Sheer, et al.
Published: (2025)
MIB: A Mechanistic Interpretability Benchmark
by: Mueller, Aaron, et al.
Published: (2025)
by: Mueller, Aaron, et al.
Published: (2025)
nnterp: A Standardized Interface for Mechanistic Interpretability of Transformers
by: Dumas, Clément
Published: (2025)
by: Dumas, Clément
Published: (2025)
Dissecting Bias in LLMs: A Mechanistic Interpretability Perspective
by: Chandna, Bhavik, et al.
Published: (2025)
by: Chandna, Bhavik, et al.
Published: (2025)
On the Mechanistic Interpretability of Neural Networks for Causality in Bio-statistics
by: Conan, Jean-Baptiste A.
Published: (2025)
by: Conan, Jean-Baptiste A.
Published: (2025)
MINAR: Mechanistic Interpretability for Neural Algorithmic Reasoning
by: He, Jesse, et al.
Published: (2026)
by: He, Jesse, et al.
Published: (2026)
Graph Neural Networks for Learning Equivariant Representations of Neural Networks
by: Kofinas, Miltiadis, et al.
Published: (2024)
by: Kofinas, Miltiadis, et al.
Published: (2024)
reward-lens: A Mechanistic Interpretability Library for Reward Models
by: Nadaf, Mohammed Suhail B
Published: (2026)
by: Nadaf, Mohammed Suhail B
Published: (2026)
A Survey on Mechanistic Interpretability for Multi-Modal Foundation Models
by: Lin, Zihao, et al.
Published: (2025)
by: Lin, Zihao, et al.
Published: (2025)
International Agreements on AI Safety: Review and Recommendations for a Conditional AI Safety Treaty
by: Scholefield, Rebecca, et al.
Published: (2025)
by: Scholefield, Rebecca, et al.
Published: (2025)
Mechanistic Interpretability of Emotion Inference in Large Language Models
by: Tak, Ala N., et al.
Published: (2025)
by: Tak, Ala N., et al.
Published: (2025)
Faithful and Stable Neuron Explanations for Trustworthy Mechanistic Interpretability
by: Yan, Ge, et al.
Published: (2025)
by: Yan, Ge, et al.
Published: (2025)
Mechanistic Interpretability for Transformer-based Time Series Classification
by: Kalnāre, Matīss, et al.
Published: (2025)
by: Kalnāre, Matīss, et al.
Published: (2025)
From Demonstrations to Rewards: Test-Time Prompt Optimization for VLM Reward Models
by: Gumbsch, Christian, et al.
Published: (2026)
by: Gumbsch, Christian, et al.
Published: (2026)
Grounding Continuous Representations in Geometry: Equivariant Neural Fields
by: Wessels, David R, et al.
Published: (2024)
by: Wessels, David R, et al.
Published: (2024)
Mechanistic Interpretability as Statistical Estimation: A Variance Analysis
by: Méloux, Maxime, et al.
Published: (2025)
by: Méloux, Maxime, et al.
Published: (2025)
Space-Time Continuous PDE Forecasting using Equivariant Neural Fields
by: Knigge, David M., et al.
Published: (2024)
by: Knigge, David M., et al.
Published: (2024)
Towards Understanding Safety Alignment: A Mechanistic Perspective from Safety Neurons
by: Chen, Jianhui, et al.
Published: (2024)
by: Chen, Jianhui, et al.
Published: (2024)
Mechanistic Interpretability with Sparse Autoencoder Neural Operators
by: Tolooshams, Bahareh, et al.
Published: (2025)
by: Tolooshams, Bahareh, et al.
Published: (2025)
Similar Items
-
Superposition as Lossy Compression: Measure with Sparse Autoencoders and Connect to Adversarial Vulnerability
by: Bereska, Leonard, et al.
Published: (2025) -
Mechanistic Neural Networks for Scientific Machine Learning
by: Pervez, Adeel, et al.
Published: (2024) -
From MLP to NeoMLP: Leveraging Self-Attention for Neural Fields
by: Kofinas, Miltiadis, et al.
Published: (2024) -
Physics-Guided Radiotherapy Treatment Planning with Deep Learning
by: Achlatis, Stefanos, et al.
Published: (2025) -
Unleashing Uncertainty: Efficient Machine Unlearning for Generative AI
by: Spartalis, Christoforos N., et al.
Published: (2025)