Stroke Lesions as a Rosetta Stone for Language Model Interpretability

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Fridriksson, Julius, Newman-Norlund, Roger D., Ahmadi, Saeed, Willis, Regan, Salman, Nadra, Warren, Kalil, Guan, Xiang, Yang, Yong, Nelakuditi, Srihari, Desai, Rutvik, Bonilha, Leonardo, Charney, Jeff, Rorden, Chris
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910011053047808
author Fridriksson, Julius
Newman-Norlund, Roger D.
Ahmadi, Saeed
Willis, Regan
Salman, Nadra
Warren, Kalil
Guan, Xiang
Yang, Yong
Nelakuditi, Srihari
Desai, Rutvik
Bonilha, Leonardo
Charney, Jeff
Rorden, Chris
author_facet Fridriksson, Julius
Newman-Norlund, Roger D.
Ahmadi, Saeed
Willis, Regan
Salman, Nadra
Warren, Kalil
Guan, Xiang
Yang, Yong
Nelakuditi, Srihari
Desai, Rutvik
Bonilha, Leonardo
Charney, Jeff
Rorden, Chris
contents Large language models (LLMs) have achieved remarkable capabilities, yet methods to verify which model components are truly necessary for language function remain limited. Current interpretability approaches rely on internal metrics and lack external validation. Here we present the Brain-LLM Unified Model (BLUM), a framework that leverages lesion-symptom mapping, the gold standard for establishing causal brain-behavior relationships for over a century, as an external reference structure for evaluating LLM perturbation effects. Using data from individuals with chronic post-stroke aphasia (N = 410), we trained symptom-to-lesion models that predict brain damage location from behavioral error profiles, applied systematic perturbations to transformer layers, administered identical clinical assessments to perturbed LLMs and human patients, and projected LLM error profiles into human lesion space. LLM error profiles were sufficiently similar to human error profiles that predicted lesions corresponded to actual lesions in error-matched humans above chance in 67% of picture naming conditions (p < 10^{-23}) and 68.3% of sentence completion conditions (p < 10^{-61}), with semantic-dominant errors mapping onto ventral-stream lesion patterns and phonemic-dominant errors onto dorsal-stream patterns. These findings open a new methodological avenue for LLM interpretability in which clinical neuroscience provides external validation, establishing human lesion-symptom mapping as a reference framework for evaluating artificial language systems and motivating direct investigation of whether behavioral alignment reflects shared computational principles.
format Preprint
id arxiv_https___arxiv_org_abs_2602_04074
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Stroke Lesions as a Rosetta Stone for Language Model Interpretability
Fridriksson, Julius
Newman-Norlund, Roger D.
Ahmadi, Saeed
Willis, Regan
Salman, Nadra
Warren, Kalil
Guan, Xiang
Yang, Yong
Nelakuditi, Srihari
Desai, Rutvik
Bonilha, Leonardo
Charney, Jeff
Rorden, Chris
Machine Learning
Computation and Language
I.2.6; I.2.7; I.6.4; J.3
Large language models (LLMs) have achieved remarkable capabilities, yet methods to verify which model components are truly necessary for language function remain limited. Current interpretability approaches rely on internal metrics and lack external validation. Here we present the Brain-LLM Unified Model (BLUM), a framework that leverages lesion-symptom mapping, the gold standard for establishing causal brain-behavior relationships for over a century, as an external reference structure for evaluating LLM perturbation effects. Using data from individuals with chronic post-stroke aphasia (N = 410), we trained symptom-to-lesion models that predict brain damage location from behavioral error profiles, applied systematic perturbations to transformer layers, administered identical clinical assessments to perturbed LLMs and human patients, and projected LLM error profiles into human lesion space. LLM error profiles were sufficiently similar to human error profiles that predicted lesions corresponded to actual lesions in error-matched humans above chance in 67% of picture naming conditions (p < 10^{-23}) and 68.3% of sentence completion conditions (p < 10^{-61}), with semantic-dominant errors mapping onto ventral-stream lesion patterns and phonemic-dominant errors onto dorsal-stream patterns. These findings open a new methodological avenue for LLM interpretability in which clinical neuroscience provides external validation, establishing human lesion-symptom mapping as a reference framework for evaluating artificial language systems and motivating direct investigation of whether behavioral alignment reflects shared computational principles.
title Stroke Lesions as a Rosetta Stone for Language Model Interpretability
topic Machine Learning
Computation and Language
I.2.6; I.2.7; I.6.4; J.3
url https://arxiv.org/abs/2602.04074