Mechanistic Anomaly Detection for "Quirky" Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Johnston, David O., Chakraborty, Arkajyoti, Belrose, Nora |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Eliciting Latent Knowledge from Quirky Language Models
by: Mallen, Alex, et al.
Published: (2023)
by: Mallen, Alex, et al.
Published: (2023)
Partially Rewriting a Transformer in Natural Language
by: Paulo, Gonçalo, et al.
Published: (2025)
by: Paulo, Gonçalo, et al.
Published: (2025)
Automatically Interpreting Millions of Features in Large Language Models
by: Paulo, Gonçalo, et al.
Published: (2024)
by: Paulo, Gonçalo, et al.
Published: (2024)
Refusal in LLMs is an Affine Function
by: Marshall, Thomas, et al.
Published: (2024)
by: Marshall, Thomas, et al.
Published: (2024)
Examining Two Hop Reasoning Through Information Content Scaling
by: Johnston, David, et al.
Published: (2025)
by: Johnston, David, et al.
Published: (2025)
Does Transformer Interpretability Transfer to RNNs?
by: Paulo, Gonçalo, et al.
Published: (2024)
by: Paulo, Gonçalo, et al.
Published: (2024)
When Actions Teach You to Think: Reasoning-Action Synergy via Reinforcement Learning in Conversational Agents
by: Rawat, Mrinal, et al.
Published: (2025)
by: Rawat, Mrinal, et al.
Published: (2025)
LEACE: Perfect linear concept erasure in closed form
by: Belrose, Nora, et al.
Published: (2023)
by: Belrose, Nora, et al.
Published: (2023)
Adapting Large Language Models for Parameter-Efficient Log Anomaly Detection
by: Lim, Ying Fu, et al.
Published: (2025)
by: Lim, Ying Fu, et al.
Published: (2025)
Detecting and Understanding Vulnerabilities in Language Models via Mechanistic Interpretability
by: García-Carrasco, Jorge, et al.
Published: (2024)
by: García-Carrasco, Jorge, et al.
Published: (2024)
Mechanistic Exploration of Backdoored Large Language Model Attention Patterns
by: Baker, Mohammed Abu, et al.
Published: (2025)
by: Baker, Mohammed Abu, et al.
Published: (2025)
Eliminating Position Bias of Language Models: A Mechanistic Approach
by: Wang, Ziqi, et al.
Published: (2024)
by: Wang, Ziqi, et al.
Published: (2024)
DetoxBench: Benchmarking Large Language Models for Multitask Fraud & Abuse Detection
by: Chakraborty, Joymallya, et al.
Published: (2024)
by: Chakraborty, Joymallya, et al.
Published: (2024)
Model Directions, Not Words: Mechanistic Topic Models Using Sparse Autoencoders
by: Zheng, Carolina, et al.
Published: (2025)
by: Zheng, Carolina, et al.
Published: (2025)
On the Quantization Robustness of Diffusion Language Models in Coding Benchmarks
by: Gupta, Aarav, et al.
Published: (2026)
by: Gupta, Aarav, et al.
Published: (2026)
Beyond a Single Perspective: Text Anomaly Detection with Multi-View Language Representations
by: Liu, Yixin, et al.
Published: (2026)
by: Liu, Yixin, et al.
Published: (2026)
Binary Autoencoder for Mechanistic Interpretability of Large Language Models
by: Cho, Hakaze, et al.
Published: (2025)
by: Cho, Hakaze, et al.
Published: (2025)
Estimating the Probability of Sampling a Trained Neural Network at Random
by: Scherlis, Adam, et al.
Published: (2025)
by: Scherlis, Adam, et al.
Published: (2025)
Evaluating SAE interpretability without explanations
by: Paulo, Gonçalo, et al.
Published: (2025)
by: Paulo, Gonçalo, et al.
Published: (2025)
Slowing Learning by Erasing Simple Features
by: Quirke, Lucia, et al.
Published: (2025)
by: Quirke, Lucia, et al.
Published: (2025)
Sparse Autoencoders Trained on the Same Data Learn Different Features
by: Paulo, Gonçalo, et al.
Published: (2025)
by: Paulo, Gonçalo, et al.
Published: (2025)
Converting MLPs into Polynomials in Closed Form
by: Belrose, Nora, et al.
Published: (2025)
by: Belrose, Nora, et al.
Published: (2025)
Balancing Label Quantity and Quality for Scalable Elicitation
by: Mallen, Alex, et al.
Published: (2024)
by: Mallen, Alex, et al.
Published: (2024)
Understanding Gradient Descent through the Training Jacobian
by: Belrose, Nora, et al.
Published: (2024)
by: Belrose, Nora, et al.
Published: (2024)
Persona-aware Generative Model for Code-mixed Language
by: Sengupta, Ayan, et al.
Published: (2023)
by: Sengupta, Ayan, et al.
Published: (2023)
Temporally Consistent Factuality Probing for Large Language Models
by: Bajpai, Ashutosh, et al.
Published: (2024)
by: Bajpai, Ashutosh, et al.
Published: (2024)
Towards Token-Level Text Anomaly Detection
by: Cao, Yang, et al.
Published: (2026)
by: Cao, Yang, et al.
Published: (2026)
NLP-ADBench: NLP Anomaly Detection Benchmark
by: Li, Yuangang, et al.
Published: (2024)
by: Li, Yuangang, et al.
Published: (2024)
Latent Performance Profiling of Large Language Models
by: Chakraborty, Tanmoy, et al.
Published: (2026)
by: Chakraborty, Tanmoy, et al.
Published: (2026)
Mechanistic Unlearning: Robust Knowledge Unlearning and Editing via Mechanistic Localization
by: Guo, Phillip, et al.
Published: (2024)
by: Guo, Phillip, et al.
Published: (2024)
Reasoning Circuits in Language Models: A Mechanistic Interpretation of Syllogistic Inference
by: Kim, Geonhee, et al.
Published: (2024)
by: Kim, Geonhee, et al.
Published: (2024)
From Zero to Hero: Cold-Start Anomaly Detection
by: Reiss, Tal, et al.
Published: (2024)
by: Reiss, Tal, et al.
Published: (2024)
Multilingual Language Models Encode Script Over Linguistic Structure
by: Verma, Aastha A K, et al.
Published: (2026)
by: Verma, Aastha A K, et al.
Published: (2026)
Detecting Hope Across Languages: Multiclass Classification for Positive Online Discourse
by: Abiola, T. O., et al.
Published: (2025)
by: Abiola, T. O., et al.
Published: (2025)
Factual Knowledge in Language Models: Robustness and Anomalies under Simple Temporal Context Variations
by: Khodja, Hichem Ammar, et al.
Published: (2025)
by: Khodja, Hichem Ammar, et al.
Published: (2025)
Multilingual Needle in a Haystack: Investigating Long-Context Behavior of Multilingual Large Language Models
by: Hengle, Amey, et al.
Published: (2024)
by: Hengle, Amey, et al.
Published: (2024)
LOLAMEME: Logic, Language, Memory, Mechanistic Framework
by: Desai, Jay, et al.
Published: (2024)
by: Desai, Jay, et al.
Published: (2024)
Can Multimodal LLMs Perform Time Series Anomaly Detection?
by: Xu, Xiongxiao, et al.
Published: (2025)
by: Xu, Xiongxiao, et al.
Published: (2025)
Steering Without Breaking: Mechanistically Informed Interventions for Discrete Diffusion Language Models
by: Zhou, Hanhan, et al.
Published: (2026)
by: Zhou, Hanhan, et al.
Published: (2026)
Where Does Toxicity Live? Mechanistic Localization and Targeted Suppression in Language Models
by: Beniwal, Himanshu, et al.
Published: (2026)
by: Beniwal, Himanshu, et al.
Published: (2026)
Similar Items
-
Eliciting Latent Knowledge from Quirky Language Models
by: Mallen, Alex, et al.
Published: (2023) -
Partially Rewriting a Transformer in Natural Language
by: Paulo, Gonçalo, et al.
Published: (2025) -
Automatically Interpreting Millions of Features in Large Language Models
by: Paulo, Gonçalo, et al.
Published: (2024) -
Refusal in LLMs is an Affine Function
by: Marshall, Thomas, et al.
Published: (2024) -
Examining Two Hop Reasoning Through Information Content Scaling
by: Johnston, David, et al.
Published: (2025)