Atlas-Alignment: Making Interpretability Transferable Across Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Puri, Bruno, Berend, Jim, Lapuschkin, Sebastian, Samek, Wojciech
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911620542758912
author Puri, Bruno
Berend, Jim
Lapuschkin, Sebastian
Samek, Wojciech
author_facet Puri, Bruno
Berend, Jim
Lapuschkin, Sebastian
Samek, Wojciech
contents Interpretability is crucial for building safe, reliable, and controllable language models, yet existing interpretability pipelines remain costly and difficult to scale. Interpreting a new model typically requires training model-specific components (e.g., sparse autoencoders), followed by manual or semi-automated labeling and validation, imposing a growing "transparency tax" that does not scale with the pace of model development. We introduce Atlas-Alignment, a framework that avoids this cost by aligning the latent space of a new model to a pre-existing, labeled Concept Atlas using only shared inputs and lightweight representational alignment methods. Through quantitative and qualitative evaluations, we show that simple alignment methods enable robust semantic retrieval and steerable generation without the need for labeled concept datasets. Atlas-Alignment thus amortizes the cost of explainable AI and mechanistic interpretability: by investing in a single high-quality Concept Atlas, we can make many new models transparent and controllable at minimal marginal cost.
format Preprint
id arxiv_https___arxiv_org_abs_2510_27413
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Atlas-Alignment: Making Interpretability Transferable Across Language Models
Puri, Bruno
Berend, Jim
Lapuschkin, Sebastian
Samek, Wojciech
Machine Learning
Artificial Intelligence
Computation and Language
Interpretability is crucial for building safe, reliable, and controllable language models, yet existing interpretability pipelines remain costly and difficult to scale. Interpreting a new model typically requires training model-specific components (e.g., sparse autoencoders), followed by manual or semi-automated labeling and validation, imposing a growing "transparency tax" that does not scale with the pace of model development. We introduce Atlas-Alignment, a framework that avoids this cost by aligning the latent space of a new model to a pre-existing, labeled Concept Atlas using only shared inputs and lightweight representational alignment methods. Through quantitative and qualitative evaluations, we show that simple alignment methods enable robust semantic retrieval and steerable generation without the need for labeled concept datasets. Atlas-Alignment thus amortizes the cost of explainable AI and mechanistic interpretability: by investing in a single high-quality Concept Atlas, we can make many new models transparent and controllable at minimal marginal cost.
title Atlas-Alignment: Making Interpretability Transferable Across Language Models
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2510.27413