Mechanistic Anomaly Detection for "Quirky" Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Johnston, David O., Chakraborty, Arkajyoti, Belrose, Nora
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912321878622208
author Johnston, David O.
Chakraborty, Arkajyoti
Belrose, Nora
author_facet Johnston, David O.
Chakraborty, Arkajyoti
Belrose, Nora
contents As LLMs grow in capability, the task of supervising LLMs becomes more challenging. Supervision failures can occur if LLMs are sensitive to factors that supervisors are unaware of. We investigate Mechanistic Anomaly Detection (MAD) as a technique to augment supervision of capable models; we use internal model features to identify anomalous training signals so they can be investigated or discarded. We train detectors to flag points from the test environment that differ substantially from the training environment, and experiment with a large variety of detector features and scoring rules to detect anomalies in a set of ``quirky'' language models. We find that detectors can achieve high discrimination on some tasks, but no detector is effective across all models and tasks. MAD techniques may be effective in low-stakes applications, but advances in both detection and evaluation are likely needed if they are to be used in high stakes settings.
format Preprint
id arxiv_https___arxiv_org_abs_2504_08812
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Mechanistic Anomaly Detection for "Quirky" Language Models
Johnston, David O.
Chakraborty, Arkajyoti
Belrose, Nora
Machine Learning
Computation and Language
As LLMs grow in capability, the task of supervising LLMs becomes more challenging. Supervision failures can occur if LLMs are sensitive to factors that supervisors are unaware of. We investigate Mechanistic Anomaly Detection (MAD) as a technique to augment supervision of capable models; we use internal model features to identify anomalous training signals so they can be investigated or discarded. We train detectors to flag points from the test environment that differ substantially from the training environment, and experiment with a large variety of detector features and scoring rules to detect anomalies in a set of ``quirky'' language models. We find that detectors can achieve high discrimination on some tasks, but no detector is effective across all models and tasks. MAD techniques may be effective in low-stakes applications, but advances in both detection and evaluation are likely needed if they are to be used in high stakes settings.
title Mechanistic Anomaly Detection for "Quirky" Language Models
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2504.08812