Mechanistic Interpretability Needs Philosophy

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Williams, Iwan, Oldenburg, Ninell, Dhar, Ruchira, Hatherley, Joshua, Fierro, Constanza, Rajcic, Nina, Schiller, Sandrine R., Stamatiou, Filippos, Søgaard, Anders
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916026386481152
author Williams, Iwan
Oldenburg, Ninell
Dhar, Ruchira
Hatherley, Joshua
Fierro, Constanza
Rajcic, Nina
Schiller, Sandrine R.
Stamatiou, Filippos
Søgaard, Anders
author_facet Williams, Iwan
Oldenburg, Ninell
Dhar, Ruchira
Hatherley, Joshua
Fierro, Constanza
Rajcic, Nina
Schiller, Sandrine R.
Stamatiou, Filippos
Søgaard, Anders
contents Mechanistic interpretability (MI) aims to explain how neural networks work by uncovering their underlying mechanisms. As the field grows in influence, it is increasingly important to examine not just models themselves, but the assumptions, concepts and explanatory strategies implicit in MI research. We argue that mechanistic interpretability needs philosophy as an ongoing partner in clarifying its concepts, refining its methods, and navigating the epistemic and ethical complexities of interpreting AI systems. There is significant unrealised potential for progress in MI to be gained through deeper engagement with philosophers and philosophical frameworks. Taking three open problems from the MI literature as examples, this paper illustrates the value philosophy can add to MI research, and outlines a path toward deeper interdisciplinary dialogue.
format Preprint
id arxiv_https___arxiv_org_abs_2506_18852
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Mechanistic Interpretability Needs Philosophy
Williams, Iwan
Oldenburg, Ninell
Dhar, Ruchira
Hatherley, Joshua
Fierro, Constanza
Rajcic, Nina
Schiller, Sandrine R.
Stamatiou, Filippos
Søgaard, Anders
Computation and Language
Artificial Intelligence
Mechanistic interpretability (MI) aims to explain how neural networks work by uncovering their underlying mechanisms. As the field grows in influence, it is increasingly important to examine not just models themselves, but the assumptions, concepts and explanatory strategies implicit in MI research. We argue that mechanistic interpretability needs philosophy as an ongoing partner in clarifying its concepts, refining its methods, and navigating the epistemic and ethical complexities of interpreting AI systems. There is significant unrealised potential for progress in MI to be gained through deeper engagement with philosophers and philosophical frameworks. Taking three open problems from the MI literature as examples, this paper illustrates the value philosophy can add to MI research, and outlines a path toward deeper interdisciplinary dialogue.
title Mechanistic Interpretability Needs Philosophy
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2506.18852