Sparse Autoencoders Do Not Find Canonical Units of Analysis
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Leask, Patrick, Bussmann, Bart, Pearce, Michael, Bloom, Joseph, Tigges, Curt, Moubayed, Noura Al, Sharkey, Lee, Nanda, Neel |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
BatchTopK Sparse Autoencoders
von: Bussmann, Bart, et al.
Veröffentlicht: (2024)
von: Bussmann, Bart, et al.
Veröffentlicht: (2024)
Inference-Time Decomposition of Activations (ITDA): A Scalable Approach to Interpreting Large Language Models
von: Leask, Patrick, et al.
Veröffentlicht: (2025)
von: Leask, Patrick, et al.
Veröffentlicht: (2025)
Learning Multi-Level Features with Matryoshka Sparse Autoencoders
von: Bussmann, Bart, et al.
Veröffentlicht: (2025)
von: Bussmann, Bart, et al.
Veröffentlicht: (2025)
LLMs for LLMs: A Structured Prompting Methodology for Long Legal Documents
von: Klem, Strahinja, et al.
Veröffentlicht: (2025)
von: Klem, Strahinja, et al.
Veröffentlicht: (2025)
MuLD: The Multitask Long Document Benchmark
von: Hudson, G Thomas, et al.
Veröffentlicht: (2022)
von: Hudson, G Thomas, et al.
Veröffentlicht: (2022)
Early Detection and Reduction of Memorisation for Domain Adaptation and Instruction Tuning
von: Slack, Dean L., et al.
Veröffentlicht: (2025)
von: Slack, Dean L., et al.
Veröffentlicht: (2025)
Interpreting Attention Layer Outputs with Sparse Autoencoders
von: Kissane, Connor, et al.
Veröffentlicht: (2024)
von: Kissane, Connor, et al.
Veröffentlicht: (2024)
Interpretable Embeddings with Sparse Autoencoders: A Data Analysis Toolkit
von: Jiang, Nick, et al.
Veröffentlicht: (2025)
von: Jiang, Nick, et al.
Veröffentlicht: (2025)
Are Sparse Autoencoders Useful? A Case Study in Sparse Probing
von: Kantamneni, Subhash, et al.
Veröffentlicht: (2025)
von: Kantamneni, Subhash, et al.
Veröffentlicht: (2025)
Minimal and Mechanistic Conditions for Behavioral Self-Awareness in LLMs
von: Bozoukov, Matthew, et al.
Veröffentlicht: (2025)
von: Bozoukov, Matthew, et al.
Veröffentlicht: (2025)
SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability
von: Karvonen, Adam, et al.
Veröffentlicht: (2025)
von: Karvonen, Adam, et al.
Veröffentlicht: (2025)
Breaking Down Financial News Impact: A Novel AI Approach with Geometric Hypergraphs
von: Harit, Anoushka, et al.
Veröffentlicht: (2024)
von: Harit, Anoushka, et al.
Veröffentlicht: (2024)
Improving Dictionary Learning with Gated Sparse Autoencoders
von: Rajamanoharan, Senthooran, et al.
Veröffentlicht: (2024)
von: Rajamanoharan, Senthooran, et al.
Veröffentlicht: (2024)
Interpretability as Compression: Reconsidering SAE Explanations of Neural Activations with MDL-SAEs
von: Ayonrinde, Kola, et al.
Veröffentlicht: (2024)
von: Ayonrinde, Kola, et al.
Veröffentlicht: (2024)
How Well Do Models Follow Their Constitutions?
von: Jakkli, Arya, et al.
Veröffentlicht: (2026)
von: Jakkli, Arya, et al.
Veröffentlicht: (2026)
Towards Principled Evaluations of Sparse Autoencoders for Interpretability and Control
von: Makelov, Aleksandar, et al.
Veröffentlicht: (2024)
von: Makelov, Aleksandar, et al.
Veröffentlicht: (2024)
A is for Absorption: Studying Feature Splitting and Absorption in Sparse Autoencoders
von: Chanin, David, et al.
Veröffentlicht: (2024)
von: Chanin, David, et al.
Veröffentlicht: (2024)
Evaluating Sparse Autoencoders on Targeted Concept Erasure Tasks
von: Karvonen, Adam, et al.
Veröffentlicht: (2024)
von: Karvonen, Adam, et al.
Veröffentlicht: (2024)
Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2
von: Lieberum, Tom, et al.
Veröffentlicht: (2024)
von: Lieberum, Tom, et al.
Veröffentlicht: (2024)
Finding Belief Geometries with Sparse Autoencoders
von: Levinson, Matthew
Veröffentlicht: (2026)
von: Levinson, Matthew
Veröffentlicht: (2026)
AttenCraft: Attention-guided Disentanglement of Multiple Concepts for Text-to-Image Customization
von: Shentu, Junjie, et al.
Veröffentlicht: (2024)
von: Shentu, Junjie, et al.
Veröffentlicht: (2024)
Textual Localization: Decomposing Multi-concept Images for Subject-Driven Text-to-Image Generation
von: Shentu, Junjie, et al.
Veröffentlicht: (2024)
von: Shentu, Junjie, et al.
Veröffentlicht: (2024)
Audio Contrastive-based Fine-tuning: Decoupling Representation Learning and Classification
von: Wang, Yang, et al.
Veröffentlicht: (2023)
von: Wang, Yang, et al.
Veröffentlicht: (2023)
Do I Know This Entity? Knowledge Awareness and Hallucinations in Language Models
von: Ferrando, Javier, et al.
Veröffentlicht: (2024)
von: Ferrando, Javier, et al.
Veröffentlicht: (2024)
ReSAE: Residualized Sparse Autoencoders for Multi-Layer Transformer Interventions
von: Poduval, Prathyush, et al.
Veröffentlicht: (2026)
von: Poduval, Prathyush, et al.
Veröffentlicht: (2026)
Identifying Sparsely Active Circuits Through Local Loss Landscape Decomposition
von: Chrisman, Brianna, et al.
Veröffentlicht: (2025)
von: Chrisman, Brianna, et al.
Veröffentlicht: (2025)
Explorations of Self-Repair in Language Models
von: Rushing, Cody, et al.
Veröffentlicht: (2024)
von: Rushing, Cody, et al.
Veröffentlicht: (2024)
Towards Best Practices of Activation Patching in Language Models: Metrics and Methods
von: Zhang, Fred, et al.
Veröffentlicht: (2023)
von: Zhang, Fred, et al.
Veröffentlicht: (2023)
Do Sparse Autoencoders Capture Concept Manifolds?
von: Bhalla, Usha, et al.
Veröffentlicht: (2026)
von: Bhalla, Usha, et al.
Veröffentlicht: (2026)
RAR-b: Reasoning as Retrieval Benchmark
von: Xiao, Chenghao, et al.
Veröffentlicht: (2024)
von: Xiao, Chenghao, et al.
Veröffentlicht: (2024)
Identifying Functionally Important Features with End-to-End Sparse Dictionary Learning
von: Braun, Dan, et al.
Veröffentlicht: (2024)
von: Braun, Dan, et al.
Veröffentlicht: (2024)
Transformer-Based Models Are Not Yet Perfect At Learning to Emulate Structural Recursion
von: Zhang, Dylan, et al.
Veröffentlicht: (2024)
von: Zhang, Dylan, et al.
Veröffentlicht: (2024)
Bidirectional Variational Autoencoders
von: Kosko, Bart, et al.
Veröffentlicht: (2025)
von: Kosko, Bart, et al.
Veröffentlicht: (2025)
LLM Circuit Analyses Are Consistent Across Training and Scale
von: Tigges, Curt, et al.
Veröffentlicht: (2024)
von: Tigges, Curt, et al.
Veröffentlicht: (2024)
Transcoders Find Interpretable LLM Feature Circuits
von: Dunefsky, Jacob, et al.
Veröffentlicht: (2024)
von: Dunefsky, Jacob, et al.
Veröffentlicht: (2024)
Investigating Permutation-Invariant Discrete Representation Learning for Spatially Aligned Images
von: Stirling, Jamie S. J., et al.
Veröffentlicht: (2026)
von: Stirling, Jamie S. J., et al.
Veröffentlicht: (2026)
Emergent Misalignment is Easy, Narrow Misalignment is Hard
von: Soligo, Anna, et al.
Veröffentlicht: (2026)
von: Soligo, Anna, et al.
Veröffentlicht: (2026)
Convergent Linear Representations of Emergent Misalignment
von: Soligo, Anna, et al.
Veröffentlicht: (2025)
von: Soligo, Anna, et al.
Veröffentlicht: (2025)
SAFE: Finding Sparse and Flat Minima to Improve Pruning
von: Lee, Dongyeop, et al.
Veröffentlicht: (2025)
von: Lee, Dongyeop, et al.
Veröffentlicht: (2025)
Sparse Autoencoders, Again?
von: Lu, Yin, et al.
Veröffentlicht: (2025)
von: Lu, Yin, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
BatchTopK Sparse Autoencoders
von: Bussmann, Bart, et al.
Veröffentlicht: (2024) -
Inference-Time Decomposition of Activations (ITDA): A Scalable Approach to Interpreting Large Language Models
von: Leask, Patrick, et al.
Veröffentlicht: (2025) -
Learning Multi-Level Features with Matryoshka Sparse Autoencoders
von: Bussmann, Bart, et al.
Veröffentlicht: (2025) -
LLMs for LLMs: A Structured Prompting Methodology for Long Legal Documents
von: Klem, Strahinja, et al.
Veröffentlicht: (2025) -
MuLD: The Multitask Long Document Benchmark
von: Hudson, G Thomas, et al.
Veröffentlicht: (2022)