Findings of the BlackboxNLP 2025 Shared Task: Localizing Circuits and Causal Variables in Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Arad, Dana, Belinkov, Yonatan, Chen, Hanjie, Kim, Najoung, Mohebbi, Hosein, Mueller, Aaron, Sarti, Gabriele, Tutek, Martin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911283149799424
author Arad, Dana
Belinkov, Yonatan
Chen, Hanjie
Kim, Najoung
Mohebbi, Hosein
Mueller, Aaron
Sarti, Gabriele
Tutek, Martin
author_facet Arad, Dana
Belinkov, Yonatan
Chen, Hanjie
Kim, Najoung
Mohebbi, Hosein
Mueller, Aaron
Sarti, Gabriele
Tutek, Martin
contents Mechanistic interpretability (MI) seeks to uncover how language models (LMs) implement specific behaviors, yet measuring progress in MI remains challenging. The recently released Mechanistic Interpretability Benchmark (MIB; Mueller et al., 2025) provides a standardized framework for evaluating circuit and causal variable localization. Building on this foundation, the BlackboxNLP 2025 Shared Task extends MIB into a community-wide reproducible comparison of MI techniques. The shared task features two tracks: circuit localization, which assesses methods that identify causally influential components and interactions driving model behavior, and causal variable localization, which evaluates approaches that map activations into interpretable features. With three teams spanning eight different methods, participants achieved notable gains in circuit localization using ensemble and regularization strategies for circuit discovery. With one team spanning two methods, participants achieved significant gains in causal variable localization using low-dimensional and non-linear projections to featurize activation vectors. The MIB leaderboard remains open; we encourage continued work in this standard evaluation framework to measure progress in MI research going forward.
format Preprint
id arxiv_https___arxiv_org_abs_2511_18409
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Findings of the BlackboxNLP 2025 Shared Task: Localizing Circuits and Causal Variables in Language Models
Arad, Dana
Belinkov, Yonatan
Chen, Hanjie
Kim, Najoung
Mohebbi, Hosein
Mueller, Aaron
Sarti, Gabriele
Tutek, Martin
Computation and Language
Artificial Intelligence
Mechanistic interpretability (MI) seeks to uncover how language models (LMs) implement specific behaviors, yet measuring progress in MI remains challenging. The recently released Mechanistic Interpretability Benchmark (MIB; Mueller et al., 2025) provides a standardized framework for evaluating circuit and causal variable localization. Building on this foundation, the BlackboxNLP 2025 Shared Task extends MIB into a community-wide reproducible comparison of MI techniques. The shared task features two tracks: circuit localization, which assesses methods that identify causally influential components and interactions driving model behavior, and causal variable localization, which evaluates approaches that map activations into interpretable features. With three teams spanning eight different methods, participants achieved notable gains in circuit localization using ensemble and regularization strategies for circuit discovery. With one team spanning two methods, participants achieved significant gains in causal variable localization using low-dimensional and non-linear projections to featurize activation vectors. The MIB leaderboard remains open; we encourage continued work in this standard evaluation framework to measure progress in MI research going forward.
title Findings of the BlackboxNLP 2025 Shared Task: Localizing Circuits and Causal Variables in Language Models
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2511.18409