Mechanistic understanding and validation of large AI models with SemanticLens

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Dreyer, Maximilian, Berend, Jim, Labarta, Tobias, Vielhaben, Johanna, Wiegand, Thomas, Lapuschkin, Sebastian, Samek, Wojciech
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916559296921600
author Dreyer, Maximilian
Berend, Jim
Labarta, Tobias
Vielhaben, Johanna
Wiegand, Thomas
Lapuschkin, Sebastian
Samek, Wojciech
author_facet Dreyer, Maximilian
Berend, Jim
Labarta, Tobias
Vielhaben, Johanna
Wiegand, Thomas
Lapuschkin, Sebastian
Samek, Wojciech
contents Unlike human-engineered systems such as aeroplanes, where each component's role and dependencies are well understood, the inner workings of AI models remain largely opaque, hindering verifiability and undermining trust. This paper introduces SemanticLens, a universal explanation method for neural networks that maps hidden knowledge encoded by components (e.g., individual neurons) into the semantically structured, multimodal space of a foundation model such as CLIP. In this space, unique operations become possible, including (i) textual search to identify neurons encoding specific concepts, (ii) systematic analysis and comparison of model representations, (iii) automated labelling of neurons and explanation of their functional roles, and (iv) audits to validate decision-making against requirements. Fully scalable and operating without human input, SemanticLens is shown to be effective for debugging and validation, summarizing model knowledge, aligning reasoning with expectations (e.g., adherence to the ABCDE-rule in melanoma classification), and detecting components tied to spurious correlations and their associated training data. By enabling component-level understanding and validation, the proposed approach helps bridge the "trust gap" between AI models and traditional engineered systems. We provide code for SemanticLens on https://github.com/jim-berend/semanticlens and a demo on https://semanticlens.hhi-research-insights.eu.
format Preprint
id arxiv_https___arxiv_org_abs_2501_05398
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Mechanistic understanding and validation of large AI models with SemanticLens
Dreyer, Maximilian
Berend, Jim
Labarta, Tobias
Vielhaben, Johanna
Wiegand, Thomas
Lapuschkin, Sebastian
Samek, Wojciech
Machine Learning
Artificial Intelligence
Unlike human-engineered systems such as aeroplanes, where each component's role and dependencies are well understood, the inner workings of AI models remain largely opaque, hindering verifiability and undermining trust. This paper introduces SemanticLens, a universal explanation method for neural networks that maps hidden knowledge encoded by components (e.g., individual neurons) into the semantically structured, multimodal space of a foundation model such as CLIP. In this space, unique operations become possible, including (i) textual search to identify neurons encoding specific concepts, (ii) systematic analysis and comparison of model representations, (iii) automated labelling of neurons and explanation of their functional roles, and (iv) audits to validate decision-making against requirements. Fully scalable and operating without human input, SemanticLens is shown to be effective for debugging and validation, summarizing model knowledge, aligning reasoning with expectations (e.g., adherence to the ABCDE-rule in melanoma classification), and detecting components tied to spurious correlations and their associated training data. By enabling component-level understanding and validation, the proposed approach helps bridge the "trust gap" between AI models and traditional engineered systems. We provide code for SemanticLens on https://github.com/jim-berend/semanticlens and a demo on https://semanticlens.hhi-research-insights.eu.
title Mechanistic understanding and validation of large AI models with SemanticLens
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2501.05398