Evaluating Contrast Localizer for Identifying Causal Units in Social & Mathematical Tasks in Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jamaa, Yassine, AlKhamissi, Badr, Ghosh, Satrajit, Schrimpf, Martin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911118394392576
author Jamaa, Yassine
AlKhamissi, Badr
Ghosh, Satrajit
Schrimpf, Martin
author_facet Jamaa, Yassine
AlKhamissi, Badr
Ghosh, Satrajit
Schrimpf, Martin
contents This work adapts a neuroscientific contrast localizer to pinpoint causally relevant units for Theory of Mind (ToM) and mathematical reasoning tasks in large language models (LLMs) and vision-language models (VLMs). Across 11 LLMs and 5 VLMs ranging in size from 3B to 90B parameters, we localize top-activated units using contrastive stimulus sets and assess their causal role via targeted ablations. We compare the effect of lesioning functionally selected units against low-activation and randomly selected units on downstream accuracy across established ToM and mathematical benchmarks. Contrary to expectations, low-activation units sometimes produced larger performance drops than the highly activated ones, and units derived from the mathematical localizer often impaired ToM performance more than those from the ToM localizer. These findings call into question the causal relevance of contrast-based localizers and highlight the need for broader stimulus sets and more accurately capture task-specific units.
format Preprint
id arxiv_https___arxiv_org_abs_2508_08276
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Evaluating Contrast Localizer for Identifying Causal Units in Social & Mathematical Tasks in Language Models
Jamaa, Yassine
AlKhamissi, Badr
Ghosh, Satrajit
Schrimpf, Martin
Computation and Language
Artificial Intelligence
This work adapts a neuroscientific contrast localizer to pinpoint causally relevant units for Theory of Mind (ToM) and mathematical reasoning tasks in large language models (LLMs) and vision-language models (VLMs). Across 11 LLMs and 5 VLMs ranging in size from 3B to 90B parameters, we localize top-activated units using contrastive stimulus sets and assess their causal role via targeted ablations. We compare the effect of lesioning functionally selected units against low-activation and randomly selected units on downstream accuracy across established ToM and mathematical benchmarks. Contrary to expectations, low-activation units sometimes produced larger performance drops than the highly activated ones, and units derived from the mathematical localizer often impaired ToM performance more than those from the ToM localizer. These findings call into question the causal relevance of contrast-based localizers and highlight the need for broader stimulus sets and more accurately capture task-specific units.
title Evaluating Contrast Localizer for Identifying Causal Units in Social & Mathematical Tasks in Language Models
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2508.08276