Mechanistic Interpretability Needs Philosophy
Fuente:
arXiv
Saved in:
| Main Authors: | Williams, Iwan, Oldenburg, Ninell, Dhar, Ruchira, Hatherley, Joshua, Fierro, Constanza, Rajcic, Nina, Schiller, Sandrine R., Stamatiou, Filippos, Søgaard, Anders |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Postmortem avatars in grief therapy: Prospects, ethics, and governance
by: Hatherley, Joshua, et al.
Published: (2026)
by: Hatherley, Joshua, et al.
Published: (2026)
Realist and Pluralist Conceptions of Intelligence and Their Implications on AI Research
by: Oldenburg, Ninell, et al.
Published: (2025)
by: Oldenburg, Ninell, et al.
Published: (2025)
Defining Knowledge: Bridging Epistemology and Large Language Models
by: Fierro, Constanza, et al.
Published: (2024)
by: Fierro, Constanza, et al.
Published: (2024)
Beyond Technocratic XAI: The Who, What & How in Explanation Design
by: Dhar, Ruchira, et al.
Published: (2025)
by: Dhar, Ruchira, et al.
Published: (2025)
On the Measure of a Model: From Intelligence to Generality
by: Dhar, Ruchira, et al.
Published: (2025)
by: Dhar, Ruchira, et al.
Published: (2025)
From Words to Worlds: Compositionality for Cognitive Architectures
by: Dhar, Ruchira, et al.
Published: (2024)
by: Dhar, Ruchira, et al.
Published: (2024)
Ethical Concern Identification in NLP: A Corpus of ACL Anthology Ethics Statements
by: Karamolegkou, Antonia, et al.
Published: (2024)
by: Karamolegkou, Antonia, et al.
Published: (2024)
Evaluating Adjective-Noun Compositionality in LLMs: Functional vs Representational Perspectives
by: Dhar, Ruchira, et al.
Published: (2026)
by: Dhar, Ruchira, et al.
Published: (2026)
Evaluation Revisited: A Taxonomy of Evaluation Concerns in Natural Language Processing
by: Dhar, Ruchira, et al.
Published: (2026)
by: Dhar, Ruchira, et al.
Published: (2026)
Goal-Directedness is in the Eye of the Beholder
by: Rajcic, Nina, et al.
Published: (2025)
by: Rajcic, Nina, et al.
Published: (2025)
The Intercepted Self: How Generative AI Challenges the Dynamics of the Relational Self
by: Schiller, Sandrine R., et al.
Published: (2025)
by: Schiller, Sandrine R., et al.
Published: (2025)
Factual Consistency of Multilingual Pretrained Language Models
by: Fierro, Constanza, et al.
Published: (2022)
by: Fierro, Constanza, et al.
Published: (2022)
Do Vision and Language Models Share Concepts? A Vector Space Alignment Study
by: Li, Jiaang, et al.
Published: (2023)
by: Li, Jiaang, et al.
Published: (2023)
Does Instruction Tuning Make LLMs More Consistent?
by: Fierro, Constanza, et al.
Published: (2024)
by: Fierro, Constanza, et al.
Published: (2024)
The Stories We Govern By: AI, Risk, and the Power of Imaginaries
by: Oldenburg, Ninell, et al.
Published: (2025)
by: Oldenburg, Ninell, et al.
Published: (2025)
Federated learning, ethics, and the double black box problem in medical AI
by: Hatherley, Joshua, et al.
Published: (2025)
by: Hatherley, Joshua, et al.
Published: (2025)
How Do Multilingual Language Models Remember Facts?
by: Fierro, Constanza, et al.
Published: (2024)
by: Fierro, Constanza, et al.
Published: (2024)
Can structural correspondences ground real world representational content in Large Language Models?
by: Williams, Iwan
Published: (2025)
by: Williams, Iwan
Published: (2025)
Learning and Sustaining Shared Normative Systems via Bayesian Rule Induction in Markov Games
by: Oldenburg, Ninell, et al.
Published: (2024)
by: Oldenburg, Ninell, et al.
Published: (2024)
Evaluating Multimodal Language Models as Visual Assistants for Visually Impaired Users
by: Karamolegkou, Antonia, et al.
Published: (2025)
by: Karamolegkou, Antonia, et al.
Published: (2025)
MuLan: A Study of Fact Mutability in Language Models
by: Fierro, Constanza, et al.
Published: (2024)
by: Fierro, Constanza, et al.
Published: (2024)
Chatting with Bots: AI, Speech Acts, and the Edge of Assertion
by: Williams, Iwan, et al.
Published: (2024)
by: Williams, Iwan, et al.
Published: (2024)
Real-Time Progress Prediction in Reasoning Language Models
by: Raaschou-Jensen, Hans Peter Lyngsøe, et al.
Published: (2025)
by: Raaschou-Jensen, Hans Peter Lyngsøe, et al.
Published: (2025)
Limits of trust in medical AI
by: Hatherley, Joshua
Published: (2025)
by: Hatherley, Joshua
Published: (2025)
Are clinicians ethically obligated to disclose their use of medical machine learning systems to patients?
by: Hatherley, Joshua
Published: (2025)
by: Hatherley, Joshua
Published: (2025)
A Mathematical Philosophy of Explanations in Mechanistic Interpretability -- The Strange Science Part I.i
by: Ayonrinde, Kola, et al.
Published: (2025)
by: Ayonrinde, Kola, et al.
Published: (2025)
A moving target in AI-assisted decision-making: Dataset shift, model updating, and the problem of update opacity
by: Hatherley, Joshua
Published: (2025)
by: Hatherley, Joshua
Published: (2025)
Data over dialogue: Why artificial intelligence is unlikely to humanise medicine
by: Hatherley, Joshua
Published: (2025)
by: Hatherley, Joshua
Published: (2025)
Understanding Subword Compositionality of Large Language Models
by: Peng, Qiwei, et al.
Published: (2025)
by: Peng, Qiwei, et al.
Published: (2025)
Comprehensive Reassessment of Large-Scale Evaluation Outcomes in LLMs: A Multifaceted Statistical Approach
by: Sun, Kun, et al.
Published: (2024)
by: Sun, Kun, et al.
Published: (2024)
Word Order and World Knowledge
by: Zhao, Qinghua, et al.
Published: (2024)
by: Zhao, Qinghua, et al.
Published: (2024)
EvalCards: A Framework for Standardized Evaluation Reporting
by: Dhar, Ruchira, et al.
Published: (2025)
by: Dhar, Ruchira, et al.
Published: (2025)
Mimetic Poet
by: McCormack, Jon, et al.
Published: (2024)
by: McCormack, Jon, et al.
Published: (2024)
Debiasing Multilingual LLMs in Cross-lingual Latent Space
by: Peng, Qiwei, et al.
Published: (2025)
by: Peng, Qiwei, et al.
Published: (2025)
Update Opacity: Epistemic Accessibility and Governance Under AI System Change
by: Ferrario, Andrea, et al.
Published: (2026)
by: Ferrario, Andrea, et al.
Published: (2026)
Mechanistic Interpretability of Emotion Inference in Large Language Models
by: Tak, Ala N., et al.
Published: (2025)
by: Tak, Ala N., et al.
Published: (2025)
Dissecting Bias in LLMs: A Mechanistic Interpretability Perspective
by: Chandna, Bhavik, et al.
Published: (2025)
by: Chandna, Bhavik, et al.
Published: (2025)
MIB: A Mechanistic Interpretability Benchmark
by: Mueller, Aaron, et al.
Published: (2025)
by: Mueller, Aaron, et al.
Published: (2025)
Diachronic and synchronic variation in the performance of adaptive machine learning systems: The ethical challenges
by: Hatherley, Joshua, et al.
Published: (2025)
by: Hatherley, Joshua, et al.
Published: (2025)
High hopes for "Deep Medicine"? AI, economics, and the future of care
by: Sparrow, Robert, et al.
Published: (2025)
by: Sparrow, Robert, et al.
Published: (2025)
Similar Items
-
Postmortem avatars in grief therapy: Prospects, ethics, and governance
by: Hatherley, Joshua, et al.
Published: (2026) -
Realist and Pluralist Conceptions of Intelligence and Their Implications on AI Research
by: Oldenburg, Ninell, et al.
Published: (2025) -
Defining Knowledge: Bridging Epistemology and Large Language Models
by: Fierro, Constanza, et al.
Published: (2024) -
Beyond Technocratic XAI: The Who, What & How in Explanation Design
by: Dhar, Ruchira, et al.
Published: (2025) -
On the Measure of a Model: From Intelligence to Generality
by: Dhar, Ruchira, et al.
Published: (2025)